Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:09:35.218633Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 100 of 146 outbound references and 3 inbound Pith citation observations for arXiv:2412.06061.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:09:35.218633Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:46.269305Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T03:32:01.301422Z
100 of 146 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4be60268-c99f-4ab9-8732-a5a9ea8e3676 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Introducing the next generation of claude., 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279686bc-f4f6-448a-987e-81dbb735b0cf · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On exact computation with an infinitely wide neural net
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b3dbd5-b8e7-4a3d-bb6c-a0dc8be9e70b · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Llm-based interaction for content generation: A case study on the perception of employees in an it department
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade0c026-4d71-49d6-8974-d0266e0524f2 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7e539e-5db4-4c83-84cb-0c991e19df62 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fast attention requires bounded entries
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946d25b4-c852-4266-9dc3-1c02bab19609 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond The Fine-Grained Complexity of Gradient Computation for Training Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94122f13-1096-4cd8-b070-b37bcc90197d · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Llm based generation of item-description for recommendation system
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c54e4e10-7f15-4e91-8735-c5ce0473dd69 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e6a9525-fe91-460d-852a-2af2898b3988 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Feature purification: How adversarial training performs robust deep learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ea6396-6df3-4a29-a3dd-14a6987877f0 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Physics of Language Models: Part 1, Learning Hierarchical Language Structures
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb75f36-af61-4fe9-a55c-b0f5cc45cccf · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Learning and generalization in overparameterized neural networks, going beyond two layers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3c83118-4231-4646-95c0-0a034bf76bce · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A convergence theory for deep learning via over-parameterization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86b3676-625c-43ab-9d8a-189810c55655 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the convergence rate of training recurrent neural networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d2edd1-b01c-4fb6-a225-c145ec4d7bf7 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Hierarchical attention network for multivariate time series long-term forecasting
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f300310-85ba-481c-8d22-7789f819720d · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Distribution of residual autocorrelations in autoregressive-integrated moving average time series models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84517f45-5df3-446f-acd7-c8a5819e00ca · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Training (overparametrized) neural networks in near-linear time
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d915b109-40f5-4085-abfa-35c2d5c8f6d3 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1b946f-992a-430b-b046-a505d3f4f377 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Language Models are Few-Shot Learners
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fd3485-5b1c-45d1-87f5-bc8cbb57989c · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Algorithm and Hardness for Dynamic Attention Maintenance in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4cbbe5-a6c6-433a-965c-41e0e96f4aa9 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fast gradient computation for rope attention in almost linear time
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc4095b5-ee5d-4a48-9f5f-520ac7c7ff1e · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Circuit Complexity Bounds for RoPE-based Transformer Architecture
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f0d32a-2edc-41d6-bde5-8d3630499b3f · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ff4ddc-7495-433d-8bdd-8880486a6c82 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond The computational limits of state-space models and mamba via the lens of circuit complexity
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68a0c98-31b4-4f95-b1d6-82a46425545f · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Universal Approximation of Visual Autoregressive Transformers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5125ce22-a7c5-4fde-900c-abedef89f749 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Hsr-enhanced sparse attention acceleration
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2271ee52-4266-4fda-801f-b7163c685639 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond TSMixer: An All-MLP Architecture for Time Series Forecasting
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cff61bb-db14-401a-b3e5-69459440e6e9 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Long sequence time-series forecasting with deep learning: A survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81cf4967-e4e5-4069-bb6a-4589df522009 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fine-tune Language Models to Approximate Unbiased In-context Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7b7dfc-97cb-4bca-8278-fa44ed907ef8 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond How to protect copyright data in optimization of large language models? In Proceedings of the AAAI Conference on Artificial Intelligence , pages 17871--17879, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc23ef02-e435-4524-bd3d-c6d3b73de431 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Long-term Forecasting with TiDE: Time-series Dense Encoder
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4e6482-9a68-47de-aac6-88b57d054b55 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Attention scheme inspired softmax regression, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac29ba1-0da6-4bc4-86da-9253652e7d9d · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28605276-17b9-40a3-a5ff-a24e45145430 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Tsmixer: Lightweight mlp-mixer model for multivariate time series forecasting
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45248676-4f9b-4f23-8f3d-9ad9b33c69fc · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac70c72-9f2e-4fab-994f-839965a530f1 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Exponential smoothing: The state of the art
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b80120-6971-43ed-a407-1be83e39932b · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond How to train your hippo: State space models with generalized orthogonal basis projections, 2022
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e68edea-4347-489e-b236-8a858a6b5786 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On computational limits of flowar models: Expressivity and efficiency
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495cb309-fe88-42e7-8137-073b791f230b · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Differential Privacy Mechanisms in Neural Tangent Kernel Regression
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96ab34b-7670-40ca-a477-039c14ab2a0c · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond An Over-parameterized Exponential Regression
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78edeb0b-25dc-4f1e-86c8-c732c9151eee · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06b5948-1f68-4e34-a8a6-cb341ac19617 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932b34b9-d478-4932-b510-99808b4b41ad · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Outlier-efficient hopfield layers for large transformer-based models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d92c60d-035f-4853-baeb-41d40292e2dc · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond HyperAttention: Long-context Attention in Near-Linear Time
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e1e965d-e8a8-4bdc-916b-467dbd36dcd0 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On computational limits of modern hopfield models: A fine-grained complexity analysis
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c974be-ef67-43d7-a0d6-87b05a7b4bcc · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fl-ntk: A neural tangent kernel-based framework for federated learning analysis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b471655-5a91-45d5-bd7b-a6866abc898f · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Long short-term memory
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 219df72f-c010-406f-b543-584eae2af93b · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Probability inequalities for sums of bounded random variables
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64544759-e629-416e-9cfe-83880f5f65e2 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Neural Network-Based Score Estimation in Diffusion Models: Optimization and Generalization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b16c28-37ad-4104-9068-84d1100d22a3 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee355ae-f657-496a-b880-f6910ef8bbb0 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond LoRA: Low-Rank Adaptation of Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 470b8d73-10b2-4f94-9fdf-7c7b426ad597 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Training Overparametrized Neural Networks in Sublinear Time
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e727fbf-f98f-4e48-834c-6dee3f2f25f5 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f92ccc-2621-4835-a162-85017b2d15f7 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Towards making the most of llm for translation quality estimation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050bbf77-e3ce-49ca-839c-ba6ed4a5e321 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Provably optimal memory capacity for modern hopfield models: Transformer-compatible dense associative memories as spherical codes
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d91b4d-1ef9-474a-942b-f33da8654067 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fddc4932-9a70-4138-b433-1018f5b97042 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On statistical rates and provably efficient criteria of latent diffusion transformers (dits)
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee20635a-29c0-4d65-b3ff-1070c32812c9 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bliva: A simple multimodal llm for better handling of text-rich visual questions
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 530a4829-7699-49ab-a39f-726c35cb36ae · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On sparse modern hopfield model
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d21088-e18e-4ce7-9fb9-01539f4d495d · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Neural tangent kernel: Convergence and generalization in neural networks
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 819664b4-7710-4a21-a097-e6626c574c4d · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Latent Space Theory for Emergent Abilities in Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 282ad5c5-c343-4f29-8c6d-d92f0cb81635 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Designing a neural network for forecasting financial and economic time series
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e554a06d-988b-4190-beca-6570f05ab024 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 428b239a-0dc3-495e-8577-709978bc78b3 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Ai in healthcare: time-series forecasting using statistical, neural, and ensemble architectures
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4ecfc20-1139-4b59-a44d-9de99fc4e8a1 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On Particle Methods for Parameter Estimation in State-Space Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61454908-a902-4eff-8170-ca9345e3d331 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1974103-8d09-4846-888f-a13d25aff3ed · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Circuit Complexity Bounds for Visual Autoregressive Model
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb1472a5-871d-4cc2-a392-d4fd3dafc9f3 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the computational complexity of self-attention
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9588800-6ef8-4d86-848d-5dd8aac14cc6 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Temporal fusion transformers for interpretable multi-horizon time series forecasting
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d38d89d-b6f2-420b-8118-032682f71a5e · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Short-term traffic flow forecasting: An experimental comparison of time-series analysis and supervised learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827738ef-5956-4c36-99eb-96621befef75 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adcde0ed-8b6d-4fbc-9b0b-ea015b8d0c92 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b49fd0-1f14-41ee-87a2-99a4640101d2 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Learning overparameterized neural networks via stochastic gradient descent on structured data
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 422a822d-42e9-4753-8e82-a4ab2566e54e · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Image creation based on transformer and generative adversarial networks
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f29b44e-3530-4f03-ad75-3f2b8504d060 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Research progress in attention mechanism in deep learning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f899bf06-a821-4fa8-be75-b55a83bf2dc5 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the expressive power of modern hopfield networks
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9988b933-2a9c-433b-9a32-1d444eafd866 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Neural algorithmic reasoning for hypergraphs with looped transformers
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b118bd-31db-4481-a68e-f7d001b1999f · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f736c0a-8ec8-46ed-92c5-0cb58966756c · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fine-grained attention i/o complexity: Comprehensive analysis for backward passes
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31154770-d751-4ef7-ad6e-53d8d44be0c9 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0d61b9-ce5e-4dc2-962b-fbcbf8621f34 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fb51ff-ae44-424a-92cf-5426d0060f35 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fee2727-9e40-455f-804d-f7c53c66d832 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Beyond linear approximations: A novel pruning approach for attention matrix
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f525e9d-5c8a-4273-9690-254d07651aef · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Exploring the frontiers of softmax: Provable optimization, applications in diffusion model, and beyond
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d435474-3ef0-4d9d-9416-2d31b86247b1 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Tighter Complexity Analysis of SparseGPT
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f726300-b392-46f4-9eec-c45fa3429d49 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Generative time series forecasting with diffusion, denoise, and disentanglement
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f973ac-7692-4998-a884-3a2fd062e6db · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c84fe1-c4fd-45ca-a201-361601bf9bad · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Looped relu mlps may be all you need as practical programmable computers
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40650d1b-5f5f-4a24-bf7f-cba95f7ab9ae · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Towards Infinite-Long Prefix in Transformer
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b3f52e-f86c-45f2-9d7c-6380beb93b51 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Tensor attention training: Provably efficient learning of higher-order transformers
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 855c0788-7175-40d2-b824-975e9327d8e7 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84eea693-32da-47de-89da-354d110cbc23 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Time-series forecasting with deep learning: a survey
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78839df3-926b-465d-967d-4b381d6cd6e4 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50531e12-03a2-4281-8856-4294ed8e9f20 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bounding the width of neural networks via coupled initialization a worst case analysis
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50961c4e-9dc3-45ea-a890-e77eeb9d7add · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Generative Artificial Intelligence for Software Engineering -- A Research Agenda
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5637a151-1615-4169-acf9-7d8d33cb651f · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c412c7a-5bb1-4208-9440-d0ecdd982205 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Gpt-4 technical report, 2024
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d8a6fa-a00a-4337-8c41-2717e31af7d3 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond An Empirical Study of the Non-determinism of ChatGPT in Code Generation
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b515489-1a8d-484b-9237-153d3ac5cfce · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Toward Understanding Why Adam Converges Faster Than SGD for Transformers
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec449c1e-5896-44fc-b543-d93850e9f415 · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Is Solving Graph Neural Tangent Kernel Equivalent to Training Graph Neural Network?
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16cf22a3-fe13-4a0f-b787-29d00310dc0d · outbound
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Feature programming for multivariate time series prediction
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19fec89-65f3-42bf-ba04-0b0dcdd72601 · inbound
Circuit Complexity Bounds for Visual Autoregressive Model Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc01840-9ca6-42bb-a828-392175214f38 · inbound
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be437bdc-f106-4352-adac-b851ecbde8da · inbound
Time Series Forecasting Through the Lens of Dynamics Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.