Pith. sign in

Paper Citation Record · LEDGER

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond

As of 17 August 2026, this Paper Citation Record lists 100 of 146 outbound references and 3 inbound Pith citation observations for arXiv:2412.06061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06061 v2

Coverage vector

measured 100 of 146 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:09:35.218633Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:46.269305Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T03:32:01.301422Z

Reference resolution

100 of 146 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4be60268-c99f-4ab9-8732-a5a9ea8e3676 · outbound

This paper cites Introducing the next generation of claude., 2024.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Introducing the next generation of claude., 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.511187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.511187Z digest=sha256:203fdc349f3db3cf12e8e547adeba65c80fcee2ecb83cc1fcbfa223b0ff80783

Observation 279686bc-f4f6-448a-987e-81dbb735b0cf · outbound

This paper cites On exact computation with an infinitely wide neural net.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On exact computation with an infinitely wide neural net

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.518261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.518261Z digest=sha256:595452ae9d86859c759fb80b1a08ce36aab7ac891ffb81d22821433776923660

Observation a1b3dbd5-b8e7-4a3d-bb6c-a0dc8be9e70b · outbound

This paper cites Llm-based interaction for content generation: A case study on the perception of employees in an it department.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Llm-based interaction for content generation: A case study on the perception of employees in an it department

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.523877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.523877Z digest=sha256:0e1ec8fd91ea432646425e93cd4488d89c470f92cfd0b43cfec76b8572e53c14

Observation ade0c026-4d71-49d6-8974-d0266e0524f2 · outbound

This paper cites Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.528602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.528602Z digest=sha256:3ddb58c4a5dc6e2477d86bc94b8bd53835e1c755a8865582c6888e3ccf7ab7e1

Observation 6d7e539e-5db4-4c83-84cb-0c991e19df62 · outbound

This paper cites Fast attention requires bounded entries.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fast attention requires bounded entries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.534157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.534157Z digest=sha256:deabe7019aeda26c662325c43b7e42b60297055be9cf7372972ec51771b8a052

Observation 946d25b4-c852-4266-9dc3-1c02bab19609 · outbound

This paper cites The Fine-Grained Complexity of Gradient Computation for Training Large Language Models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond The Fine-Grained Complexity of Gradient Computation for Training Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.541328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.541328Z digest=sha256:a9bcd6f3a63f03edfee6814b3db72157564af1c77587af8bd909319fbc0af346

Observation 94122f13-1096-4cd8-b070-b37bcc90197d · outbound

This paper cites Llm based generation of item-description for recommendation system.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Llm based generation of item-description for recommendation system

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.556757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.556757Z digest=sha256:f21ea420f206cd3d191f9c74ea6e4484571312de8ba9fc1643b00268efae2c73

Observation c54e4e10-7f15-4e91-8735-c5ce0473dd69 · outbound

This paper cites Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.566547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.566547Z digest=sha256:010ec058865f921be5a3bbafd80763d2ea7f6f3fa203ad1028b60e1103639187

Observation 4e6a9525-fe91-460d-852a-2af2898b3988 · outbound

This paper cites Feature purification: How adversarial training performs robust deep learning.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Feature purification: How adversarial training performs robust deep learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.577666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.577666Z digest=sha256:48fe90eb820183bc76b7b0929942f5f87b144b2c701f5702a6b5631344ef6802

Observation 05ea6396-6df3-4a29-a3dd-14a6987877f0 · outbound

This paper cites Physics of Language Models: Part 1, Learning Hierarchical Language Structures.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Physics of Language Models: Part 1, Learning Hierarchical Language Structures

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.588988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.588988Z digest=sha256:12f465ca53a3ca616dde1e1414f721f2fd477683bc7280fcafd809558589f5aa

Observation edb75f36-af61-4fe9-a55c-b0f5cc45cccf · outbound

This paper cites Learning and generalization in overparameterized neural networks, going beyond two layers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Learning and generalization in overparameterized neural networks, going beyond two layers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.597787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.597787Z digest=sha256:2bb353c6d044eec0c9456012a72706e7ae4df7aae53828f71b052e6de62f7bf9

Observation d3c83118-4231-4646-95c0-0a034bf76bce · outbound

This paper cites A convergence theory for deep learning via over-parameterization.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A convergence theory for deep learning via over-parameterization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.604056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.604056Z digest=sha256:ea6250a299a5e29c274ea348f58ccff916cf375a726654cf449a6c162449d1b5

Observation d86b3676-625c-43ab-9d8a-189810c55655 · outbound

This paper cites On the convergence rate of training recurrent neural networks.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the convergence rate of training recurrent neural networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.611451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.611451Z digest=sha256:5d67762088c641e13384b77d50937921e6361d334d657d83169410722544b8c4

Observation d5d2edd1-b01c-4fb6-a225-c145ec4d7bf7 · outbound

This paper cites Hierarchical attention network for multivariate time series long-term forecasting.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Hierarchical attention network for multivariate time series long-term forecasting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.617171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.617171Z digest=sha256:cf228c0c3ba5b7dfd1d19d0f78eb169124cb7beae16daaf853a40e608c8b61d0

Observation 1f300310-85ba-481c-8d22-7789f819720d · outbound

This paper cites Distribution of residual autocorrelations in autoregressive-integrated moving average time series models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Distribution of residual autocorrelations in autoregressive-integrated moving average time series models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.622387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.622387Z digest=sha256:72d99c9881a4b7870f5ecf004864478434a00297735c9a5d372a62cd14e6d912

Observation 84517f45-5df3-446f-acd7-c8a5819e00ca · outbound

This paper cites Training (overparametrized) neural networks in near-linear time.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Training (overparametrized) neural networks in near-linear time

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.628313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.628313Z digest=sha256:7b0c5a0fde73e60ec23b1f0240caa214845f2e791dd381f36d6b9e4692bd834f

Observation d915b109-40f5-4085-abfa-35c2d5c8f6d3 · outbound

This paper cites Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.635056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.635056Z digest=sha256:ef83389bd4d0835c58c9798bc32dc317c312a10d7502ce2c53befb3fd978e3ad

Observation 6b1b946f-992a-430b-b046-a505d3f4f377 · outbound

This paper cites Language Models are Few-Shot Learners.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Language Models are Few-Shot Learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.641847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.641847Z digest=sha256:78cb6022e7da916284cffeffe320d55b5a29088d0f84e98e8015bc204436d51c

Observation 20fd3485-5b1c-45d1-87f5-bc8cbb57989c · outbound

This paper cites Algorithm and Hardness for Dynamic Attention Maintenance in Large Language Models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Algorithm and Hardness for Dynamic Attention Maintenance in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.653536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.653536Z digest=sha256:d78e73844effdfffa00895983c4038c6dae89673a33a5ba0563a17d6bd4ad12d

Observation ea4cbbe5-a6c6-433a-965c-41e0e96f4aa9 · outbound

This paper cites Fast gradient computation for rope attention in almost linear time.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fast gradient computation for rope attention in almost linear time

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.663688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.663688Z digest=sha256:f95a730964d4767ddf6025727277b40c3035488dc07a1b620f17f7ca082b3ee9

Observation bc4095b5-ee5d-4a48-9f5f-520ac7c7ff1e · outbound

This paper cites Circuit Complexity Bounds for RoPE-based Transformer Architecture.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Circuit Complexity Bounds for RoPE-based Transformer Architecture

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.675235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.675235Z digest=sha256:5ca675e40fe0cd66c00049579d7370c896e0e0eeeb7b58dbeadb679df00df902

Observation 22f0d32a-2edc-41d6-bde5-8d3630499b3f · outbound

This paper cites Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.684662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.684662Z digest=sha256:7670007fc21af55e7eba9746f97ac2b7bd67078be357cccd8eff25e918b36ffa

Observation f0ff4ddc-7495-433d-8bdd-8880486a6c82 · outbound

This paper cites The computational limits of state-space models and mamba via the lens of circuit complexity.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond The computational limits of state-space models and mamba via the lens of circuit complexity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.694047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.694047Z digest=sha256:c8ef3e17b0b0167f09d9709a2b8c6ff94416e7e91d964cb896e1f355b6b5287a

Observation e68a0c98-31b4-4f95-b1d6-82a46425545f · outbound

This paper cites Universal Approximation of Visual Autoregressive Transformers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Universal Approximation of Visual Autoregressive Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.699529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.699529Z digest=sha256:e6e62c1c3a6b8d0653c9ad740d6e1f24fd4691a808b3ad0a6c5724d54739ceb4

Observation 5125ce22-a7c5-4fde-900c-abedef89f749 · outbound

This paper cites Hsr-enhanced sparse attention acceleration.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Hsr-enhanced sparse attention acceleration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.704753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.704753Z digest=sha256:3ec390e087f1d2d9de34e9cf04c96b8ac149aa503e0618361441b73a36c5bb1d

Observation 2271ee52-4266-4fda-801f-b7163c685639 · outbound

This paper cites TSMixer: An All-MLP Architecture for Time Series Forecasting.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond TSMixer: An All-MLP Architecture for Time Series Forecasting

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.721211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.721211Z digest=sha256:3fa446ba4956de7888043749115f62db3d818ea96611bc400e91e28ddcfb2c5b

Observation 8cff61bb-db14-401a-b3e5-69459440e6e9 · outbound

This paper cites Long sequence time-series forecasting with deep learning: A survey.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Long sequence time-series forecasting with deep learning: A survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.730024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.730024Z digest=sha256:3281e403f294b3ce6c3b70b6d7ded4d88a146d4165e35b9021c7a09780fa83bd

Observation 81cf4967-e4e5-4069-bb6a-4589df522009 · outbound

This paper cites Fine-tune Language Models to Approximate Unbiased In-context Learning.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fine-tune Language Models to Approximate Unbiased In-context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.737148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.737148Z digest=sha256:26fbd3802de1bb38b4efcbc594eca6e20a8cc81ec41eac421df742d9b3b4b713

Observation 3e7b7dfc-97cb-4bca-8278-fa44ed907ef8 · outbound

This paper cites How to protect copyright data in optimization of large language models? In Proceedings of the AAAI Conference on Artificial Intelligence , pages 17871--17879, 2024.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond How to protect copyright data in optimization of large language models? In Proceedings of the AAAI Conference on Artificial Intelligence , pages 17871--17879, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.747341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.747341Z digest=sha256:fb724ed6cbb2d1917ab1c0748c6523b44a046e5e546a7d53b0ae572bbe1bcd4d

Observation fc23ef02-e435-4524-bd3d-c6d3b73de431 · outbound

This paper cites Long-term Forecasting with TiDE: Time-series Dense Encoder.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Long-term Forecasting with TiDE: Time-series Dense Encoder

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.759938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.759938Z digest=sha256:3ce980273792b2f30e010e31061285e8c09c0125fe5a740c04383862f3d1c04d

Observation 2a4e6482-9a68-47de-aac6-88b57d054b55 · outbound

This paper cites Attention scheme inspired softmax regression, 2023.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Attention scheme inspired softmax regression, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.765636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.765636Z digest=sha256:d4853987311d9e7c51d7c854a6dd27f643251ced2f12a5b5c2e1582e6bae609d

Observation 9ac29ba1-0da6-4bc4-86da-9253652e7d9d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.770982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.770982Z digest=sha256:4e55e661c58636bb6cb373a2ef25e47cadb195141d49c4023e8ddcbf0165bb85

Observation 28605276-17b9-40a3-a5ff-a24e45145430 · outbound

This paper cites Tsmixer: Lightweight mlp-mixer model for multivariate time series forecasting.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Tsmixer: Lightweight mlp-mixer model for multivariate time series forecasting

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.777609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.777609Z digest=sha256:10f02300a808da6762c47ab28e0832fb4f2ef5e21b894ba7b81182e12d0a5b4a

Observation 45248676-4f9b-4f23-8f3d-9ad9b33c69fc · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.783765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.783765Z digest=sha256:4fb0d25af0b897f22541385cc25b8c86af31b8e690905b41cd80c1b2531ecb08

Observation 3ac70c72-9f2e-4fab-994f-839965a530f1 · outbound

This paper cites Exponential smoothing: The state of the art.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Exponential smoothing: The state of the art

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.790066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.790066Z digest=sha256:7b34e518c02693e794cbe7dd34bb3be39d120aebc26fc75050ebe263542466fa

Observation d4b80120-6971-43ed-a407-1be83e39932b · outbound

This paper cites How to train your hippo: State space models with generalized orthogonal basis projections, 2022.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond How to train your hippo: State space models with generalized orthogonal basis projections, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.795246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.795246Z digest=sha256:9eccdadbe74059579f2ff8921893a8d10de15fbd59b80cd50cfa877fad6120d7

Observation 9e68edea-4347-489e-b236-8a858a6b5786 · outbound

This paper cites On computational limits of flowar models: Expressivity and efficiency.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On computational limits of flowar models: Expressivity and efficiency

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.801184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.801184Z digest=sha256:d4b7f904a0f129b9d6dec25242d1135752f9399444d1c7f6cff735752755ab88

Observation 495cb309-fe88-42e7-8137-073b791f230b · outbound

This paper cites Differential Privacy Mechanisms in Neural Tangent Kernel Regression.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Differential Privacy Mechanisms in Neural Tangent Kernel Regression

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.807219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.807219Z digest=sha256:dd407938a6210f74b1d1dad7dd9c5601743fb4dd85b183fe091bdcd3ccb58885

Observation a96ab34b-7670-40ca-a477-039c14ab2a0c · outbound

This paper cites An Over-parameterized Exponential Regression.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond An Over-parameterized Exponential Regression

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.812967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.812967Z digest=sha256:aa62bc8e56bce28e8f8f32ed7e6042e17e916abd5465789747f52993ba394607

Observation 78edeb0b-25dc-4f1e-86c8-c732c9151eee · outbound

This paper cites A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.817896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.817896Z digest=sha256:40c755a570ab3753c9c0fba6b9f66e0223d1a7f23f8213054f09964ba6d6c9ef

Observation a06b5948-1f68-4e34-a8a6-cb341ac19617 · outbound

This paper cites How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.823091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.823091Z digest=sha256:55a88a9c373f1eaf697bbca7cdab2d5a8d306f6c5a6f2b7c25d7c2735ff370b5

Observation 932b34b9-d478-4932-b510-99808b4b41ad · outbound

This paper cites Outlier-efficient hopfield layers for large transformer-based models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Outlier-efficient hopfield layers for large transformer-based models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.834819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.834819Z digest=sha256:3544d0eb36735af75b1c787399acc775132be527e28fb1d95f055ff122cbda17

Observation 4d92c60d-035f-4853-baeb-41d40292e2dc · outbound

This paper cites HyperAttention: Long-context Attention in Near-Linear Time.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond HyperAttention: Long-context Attention in Near-Linear Time

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.840028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.840028Z digest=sha256:2104534114ce61b9522bd78acb1feea7654255554c8764cc79e818fb47273e9a

Observation 2e1e965d-e8a8-4bdc-916b-467dbd36dcd0 · outbound

This paper cites On computational limits of modern hopfield models: A fine-grained complexity analysis.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On computational limits of modern hopfield models: A fine-grained complexity analysis

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.845921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.845921Z digest=sha256:0026a476d1ca3ca4e04cc4d1352c2bcc0d2cbe6e91f77e7b0abea292f0a8ce7f

Observation 86c974be-ef67-43d7-a0d6-87b05a7b4bcc · outbound

This paper cites Fl-ntk: A neural tangent kernel-based framework for federated learning analysis.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fl-ntk: A neural tangent kernel-based framework for federated learning analysis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.852057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.852057Z digest=sha256:ee1b8726a908810d8f33fe7836ed612222d637bd13017d332b811ba49ee3cbaa

Observation 7b471655-5a91-45d5-bd7b-a6866abc898f · outbound

This paper cites Long short-term memory.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Long short-term memory

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.858339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.858339Z digest=sha256:b077b202f46fea9a8aaa3382d6c7428a1deb8ddab63b92365fb16078a99493d4

Observation 219df72f-c010-406f-b543-584eae2af93b · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Probability inequalities for sums of bounded random variables

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.864312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.864312Z digest=sha256:3d346a318a4a76984f9d138dbc4e0ea05362e0220661318d2bc8c3eebcbbc56b

Observation 64544759-e629-416e-9cfe-83880f5f65e2 · outbound

This paper cites Neural Network-Based Score Estimation in Diffusion Models: Optimization and Generalization.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Neural Network-Based Score Estimation in Diffusion Models: Optimization and Generalization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.870994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.870994Z digest=sha256:e1cb722f959d80336a3e10bf89a11c0aa1178d6203f62c649690e1d41a9ff539

Observation e6b16c28-37ad-4104-9068-84d1100d22a3 · outbound

This paper cites Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.877783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.877783Z digest=sha256:f9418835a7e8d0ed1a52f1394a0904954ff74d9bf51e7f75f2a7b94dd7068521

Observation aee355ae-f657-496a-b880-f6910ef8bbb0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond LoRA: Low-Rank Adaptation of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.889124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.889124Z digest=sha256:4c8cb7c420fa47e82ce1c4e9964c565293fabe0662b79b5b241bb0c0058a08f7

Observation 470b8d73-10b2-4f94-9fdf-7c7b426ad597 · outbound

This paper cites Training Overparametrized Neural Networks in Sublinear Time.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Training Overparametrized Neural Networks in Sublinear Time

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.897231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.897231Z digest=sha256:33d03b79e700b07d31377fde429d61241e98b62af6949d8443adc299c2b4e37a

Observation 1e727fbf-f98f-4e48-834c-6dee3f2f25f5 · outbound

This paper cites Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.901765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.901765Z digest=sha256:5d5d0b600ce11e79435bdc86aa2301fb78e1309e649aabf9f298a7324c29b837

Observation a9f92ccc-2621-4835-a162-85017b2d15f7 · outbound

This paper cites Towards making the most of llm for translation quality estimation.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Towards making the most of llm for translation quality estimation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.906533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.906533Z digest=sha256:429610fb7e179f0cce821163971256014019fb3876f0ff2e87bd6f212d65b87f

Observation 050bbf77-e3ce-49ca-839c-ba6ed4a5e321 · outbound

This paper cites Provably optimal memory capacity for modern hopfield models: Transformer-compatible dense associative memories as spherical codes.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Provably optimal memory capacity for modern hopfield models: Transformer-compatible dense associative memories as spherical codes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.912591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.912591Z digest=sha256:ddd3594819488fa7096992446e628a5474029eb5f71cb510a12a09824e1c9ddf

Observation e9d91b4d-1ef9-474a-942b-f33da8654067 · outbound

This paper cites On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.917505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.917505Z digest=sha256:2b91cdaf624329c8a29b48246f69369b9f1af7ba6d23fd611725383f61093b07

Observation fddc4932-9a70-4138-b433-1018f5b97042 · outbound

This paper cites On statistical rates and provably efficient criteria of latent diffusion transformers (dits).

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On statistical rates and provably efficient criteria of latent diffusion transformers (dits)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.923107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.923107Z digest=sha256:ce9fb67ad126e995794eeb541a14e3832cc7ecf0058e2c160f3a1df3bf568be4

Observation ee20635a-29c0-4d65-b3ff-1070c32812c9 · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bliva: A simple multimodal llm for better handling of text-rich visual questions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.929399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.929399Z digest=sha256:c2773cfcc821a7d22512bf834c79c6799606d8fcba22a64d18f4bba2effa0b8a

Observation 530a4829-7699-49ab-a39f-726c35cb36ae · outbound

This paper cites On sparse modern hopfield model.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On sparse modern hopfield model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.935453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.935453Z digest=sha256:14dd73f57b5e7545f6bf8d79362f7cd72866e4223db6f676fb30b2ca3a446701

Observation 73d21088-e18e-4ce7-9fb9-01539f4d495d · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Neural tangent kernel: Convergence and generalization in neural networks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.942938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.942938Z digest=sha256:cbff866fecf329df643383afa596562266c767b0b833c73b9807601b77dab614

Observation 819664b4-7710-4a21-a097-e6626c574c4d · outbound

This paper cites A Latent Space Theory for Emergent Abilities in Large Language Models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Latent Space Theory for Emergent Abilities in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.949524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.949524Z digest=sha256:30ada3056708552bb5f5c772795d5e613e6d10a0d74fc241f098e9b739c288e0

Observation 282ad5c5-c343-4f29-8c6d-d92f0cb81635 · outbound

This paper cites Designing a neural network for forecasting financial and economic time series.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Designing a neural network for forecasting financial and economic time series

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.955180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.955180Z digest=sha256:d1a8c5e55259b3a044143aa25ca6b60521e332818f42aec173a7363470f0967c

Observation e554a06d-988b-4190-beca-6570f05ab024 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.961315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.961315Z digest=sha256:63b6323b011bf67185b7db71c75c7c84f70750cf236e72fcdb1b1169b1655aed

Observation 428b239a-0dc3-495e-8577-709978bc78b3 · outbound

This paper cites Ai in healthcare: time-series forecasting using statistical, neural, and ensemble architectures.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Ai in healthcare: time-series forecasting using statistical, neural, and ensemble architectures

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.968750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.968750Z digest=sha256:146e92699dfa1e3a3f58e7bb81ad597b845299cfd4a5d3cec03bc612d6a1ce62

Observation b4ecfc20-1139-4b59-a44d-9de99fc4e8a1 · outbound

This paper cites On Particle Methods for Parameter Estimation in State-Space Models.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On Particle Methods for Parameter Estimation in State-Space Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:09:36.743854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:09:34.974296Z digest=sha256:9783bbad68af110a70ad2299dd5894f1428e9a5f7b97545e8d98671d3fa6bb35

Observation 61454908-a902-4eff-8170-ca9345e3d331 · outbound

This paper cites On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.980370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.980370Z digest=sha256:ed9edc0dca41eb93dad9c54a077d64b254502fd145c15e6a933d98b28ee4256f

Observation c1974103-8d09-4846-888f-a13d25aff3ed · outbound

This paper cites Circuit Complexity Bounds for Visual Autoregressive Model.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Circuit Complexity Bounds for Visual Autoregressive Model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.987140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.987140Z digest=sha256:eb78621990da8c910aedcc4cbbb85a54a26e9af5ff25e69dfe6d56fbad8c412b

Observation fb1472a5-871d-4cc2-a392-d4fd3dafc9f3 · outbound

This paper cites On the computational complexity of self-attention.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the computational complexity of self-attention

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.992215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.992215Z digest=sha256:b8b91e48042c365c631d1a35368467bedf872d02344c88dd275e322d910f4c0b

Observation e9588800-6ef8-4d86-848d-5dd8aac14cc6 · outbound

This paper cites Temporal fusion transformers for interpretable multi-horizon time series forecasting.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Temporal fusion transformers for interpretable multi-horizon time series forecasting

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:34.996940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:34.996940Z digest=sha256:f481d70f9fa99c9726fa7c642339c350a104dafa1b00c5c2506fc9ef1a2adc63

Observation 3d38d89d-b6f2-420b-8118-032682f71a5e · outbound

This paper cites Short-term traffic flow forecasting: An experimental comparison of time-series analysis and supervised learning.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Short-term traffic flow forecasting: An experimental comparison of time-series analysis and supervised learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.001253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.001253Z digest=sha256:8f7ce5e3c5d07252a857e8f6802df203d466952f48ea234575f386b70164f547

Observation 827738ef-5956-4c36-99eb-96621befef75 · outbound

This paper cites iTransformer: Inverted Transformers Are Effective for Time Series Forecasting.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond iTransformer: Inverted Transformers Are Effective for Time Series Forecasting

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.006569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.006569Z digest=sha256:5379c2c5856d2d31583d46fffa6e823f43c4574fdec7d3403403290532758ac2

Observation adcde0ed-8b6d-4fbc-9b0b-ea015b8d0c92 · outbound

This paper cites Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.012026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.012026Z digest=sha256:b1b088416be768acf06ee0802a29992c60f90c82aced2328676570e209ea1830

Observation b7b49fd0-1f14-41ee-87a2-99a4640101d2 · outbound

This paper cites Learning overparameterized neural networks via stochastic gradient descent on structured data.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Learning overparameterized neural networks via stochastic gradient descent on structured data

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.017188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.017188Z digest=sha256:a162e0386e10fd816da5cddc49ded27f0f551550d8c287a0c5745c26c86b7806

Observation 422a822d-42e9-4753-8e82-a4ab2566e54e · outbound

This paper cites Image creation based on transformer and generative adversarial networks.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Image creation based on transformer and generative adversarial networks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.023185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.023185Z digest=sha256:5b5d629b02ca6c597cdb54a576616d07e98694bf610e1bd6b4cfbeb87806b699

Observation 8f29b44e-3530-4f03-ad75-3f2b8504d060 · outbound

This paper cites Research progress in attention mechanism in deep learning.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Research progress in attention mechanism in deep learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.028119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.028119Z digest=sha256:593e76d094e53c869717e577330ca644acb217f13f1069d19c2ff7cf28c34fc7

Observation f899bf06-a821-4fa8-be75-b55a83bf2dc5 · outbound

This paper cites On the expressive power of modern hopfield networks.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the expressive power of modern hopfield networks

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.034010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.034010Z digest=sha256:a91ca7492de648fed2340d550ccd561494e9ab5e98d10d83c6ed454961b8efd9

Observation 9988b933-2a9c-433b-9a32-1d444eafd866 · outbound

This paper cites Neural algorithmic reasoning for hypergraphs with looped transformers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Neural algorithmic reasoning for hypergraphs with looped transformers

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.039403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.039403Z digest=sha256:3901b99af3bc8fc67515de5cd073303cfe7537b6ab846b8b447a49e4c33ae4a6

Observation 25b118bd-31db-4481-a68e-f7d001b1999f · outbound

This paper cites Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.045433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.045433Z digest=sha256:3e193eaac31f3521eb2296b0d5288e0737899ec33e9b05199bc0353ad3c318bc

Observation 8f736c0a-8ec8-46ed-92c5-0cb58966756c · outbound

This paper cites Fine-grained attention i/o complexity: Comprehensive analysis for backward passes.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fine-grained attention i/o complexity: Comprehensive analysis for backward passes

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.051107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.051107Z digest=sha256:63a03fb62e0b072b80b95f5821654057b3d410a78ebca77c77d6c4a6d0d2f8dc

Observation 31154770-d751-4ef7-ad6e-53d8d44be0c9 · outbound

This paper cites Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.056576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.056576Z digest=sha256:b20a4ba5b8c2e81216e2e64103a4f22b9b6250b7e5f6193ed7e9f6ac8f49a645

Observation 7b0d61b9-ce5e-4dc2-962b-fbcbf8621f34 · outbound

This paper cites Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.063521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.063521Z digest=sha256:796dafdc6bcfa23124badf72dd988cfa689a6941c68a3fef9c7d57a7d01c45b5

Observation 96fb51ff-ae44-424a-92cf-5426d0060f35 · outbound

This paper cites On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.075009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.075009Z digest=sha256:7bcdc06ba502cbf9c43615c97037cc1b7a8cd1be24dc0bb096958988d1a027bc

Observation 3fee2727-9e40-455f-804d-f7c53c66d832 · outbound

This paper cites Beyond linear approximations: A novel pruning approach for attention matrix.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Beyond linear approximations: A novel pruning approach for attention matrix

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.081152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.081152Z digest=sha256:47ebf6c5584913727be7a6e27821545b94f071e6e3da81a0d1ad2b5e97b35ba9

Observation 9f525e9d-5c8a-4273-9690-254d07651aef · outbound

This paper cites Exploring the frontiers of softmax: Provable optimization, applications in diffusion model, and beyond.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Exploring the frontiers of softmax: Provable optimization, applications in diffusion model, and beyond

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.085975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.085975Z digest=sha256:9c43f7a0cfa25bc50ac4d72b561d12d2190ca0afb6cfa3fdc4cd944625564a26

Observation 5d435474-3ef0-4d9d-9416-2d31b86247b1 · outbound

This paper cites A Tighter Complexity Analysis of SparseGPT.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Tighter Complexity Analysis of SparseGPT

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.093182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.093182Z digest=sha256:664c192e8a536a896290ce1972b7804fd463e63762012a1a7d8be28d757b5126

Observation 4f726300-b392-46f4-9eec-c45fa3429d49 · outbound

This paper cites Generative time series forecasting with diffusion, denoise, and disentanglement.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Generative time series forecasting with diffusion, denoise, and disentanglement

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.102218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.102218Z digest=sha256:07a4a94bcaebe8cf034e0d17b270c4a8de3a73a971e30ae010f3f564a15107c9

Observation a9f973ac-7692-4998-a884-3a2fd062e6db · outbound

This paper cites Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.108167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.108167Z digest=sha256:cd112e71156b4dedcb7cb9594613ce2229a20f2eb84751882928f5e5e24fcc05

Observation 91c84fe1-c4fd-45ca-a201-361601bf9bad · outbound

This paper cites Looped relu mlps may be all you need as practical programmable computers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Looped relu mlps may be all you need as practical programmable computers

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.115021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.115021Z digest=sha256:cba0d9cb554dd3d1e8ca0e50fd5512dc9a8e5291a9c1b413aa5184f5d259faaa

Observation 40650d1b-5f5f-4a24-bf7f-cba95f7ab9ae · outbound

This paper cites Towards Infinite-Long Prefix in Transformer.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Towards Infinite-Long Prefix in Transformer

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.120554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.120554Z digest=sha256:d25bc396f3979e3fd1cea8cc12e5bb1cb1dbfe672177b662a1ff58e0967917c6

Observation 32b3f52e-f86c-45f2-9d7c-6380beb93b51 · outbound

This paper cites Tensor attention training: Provably efficient learning of higher-order transformers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Tensor attention training: Provably efficient learning of higher-order transformers

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.126724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.126724Z digest=sha256:f99eaf681dff91c771465879f8fe350e38f4584ddec8ee18b7bcbac84883378d

Observation 855c0788-7175-40d2-b824-975e9327d8e7 · outbound

This paper cites Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.132606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.132606Z digest=sha256:ee1268a9806b9e06bb6d1978a58fb75b8206468439738449bee63588c97b1d6f

Observation 84eea693-32da-47de-89da-354d110cbc23 · outbound

This paper cites Time-series forecasting with deep learning: a survey.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Time-series forecasting with deep learning: a survey

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.139281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.139281Z digest=sha256:fcacff6ff1dd01ebc26b77b0871c8e3533f8541141c16ec79c8ef5dac2cbff76

Observation 78839df3-926b-465d-967d-4b381d6cd6e4 · outbound

This paper cites U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.149514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.149514Z digest=sha256:5e33b3ac11b2a55fc926cb223fb91597cdbcdb8f3464e4a2e947f18d07d734de

Observation 50531e12-03a2-4281-8856-4294ed8e9f20 · outbound

This paper cites Bounding the width of neural networks via coupled initialization a worst case analysis.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Bounding the width of neural networks via coupled initialization a worst case analysis

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.158237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.158237Z digest=sha256:bc71071b9b8c5488ff968352265ee4c20ed3384a63e807a581cf766f4afff4ad

Observation 50961c4e-9dc3-45ea-a890-e77eeb9d7add · outbound

This paper cites Generative Artificial Intelligence for Software Engineering -- A Research Agenda.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Generative Artificial Intelligence for Software Engineering -- A Research Agenda

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.166950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.166950Z digest=sha256:77f84826212ffeada90c11ad52c7c7f7b481f788f81d80e0dc5b56ac2856b261

Observation 5637a151-1615-4169-acf9-7d8d33cb651f · outbound

This paper cites A Time Series is Worth 64 Words: Long-term Forecasting with Transformers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond A Time Series is Worth 64 Words: Long-term Forecasting with Transformers

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.176976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.176976Z digest=sha256:1a5a72e7517262c47e8dfa8f023e90a23d8c787161eb6bfe7f48a4f6f227d965

Observation 9c412c7a-5bb1-4208-9440-d0ecdd982205 · outbound

This paper cites Gpt-4 technical report, 2024.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Gpt-4 technical report, 2024

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.194947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.194947Z digest=sha256:c37a29d74d00f091a886945be3afcbb370ddb8b0e3b053e887c55426af33fc9c

Observation 23d8a6fa-a00a-4337-8c41-2717e31af7d3 · outbound

This paper cites An Empirical Study of the Non-determinism of ChatGPT in Code Generation.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond An Empirical Study of the Non-determinism of ChatGPT in Code Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.201442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.201442Z digest=sha256:d3fc5cc7b76a899f46b59fca2fb84a4c9f8dc22c986b89315dc1ef4ca342ecfa

Observation 0b515489-1a8d-484b-9237-153d3ac5cfce · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.207510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.207510Z digest=sha256:46159e16e2c24212d438ad593326e65aadc7dccedad99f705f42267361f7cd40

Observation ec449c1e-5896-44fc-b543-d93850e9f415 · outbound

This paper cites Is Solving Graph Neural Tangent Kernel Equivalent to Training Graph Neural Network?.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Is Solving Graph Neural Tangent Kernel Equivalent to Training Graph Neural Network?

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.212872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.212872Z digest=sha256:4519d8d678b7b9bbdbce6740565fdacd8cdec098e9bb2b1fe878bb268a83298e

Observation 16cf22a3-fe13-4a0f-b787-29d00310dc0d · outbound

This paper cites Feature programming for multivariate time series prediction.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Feature programming for multivariate time series prediction

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.218633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.218633Z digest=sha256:787dc0a919f5f364e4b60af5b454a60596ce94a58f593577bd2e584555e865d2

Pith citing papers

Observation e19fec89-65f3-42bf-ba04-0b0dcdd72601 · inbound

Circuit Complexity Bounds for Visual Autoregressive Model cites this paper.

Circuit Complexity Bounds for Visual Autoregressive Model Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:46.269305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:46.269305Z digest=sha256:67861cc674c0e9a9397891977af9f3e3b01bab3e35d9c9684d28eec826a0ed09

Observation cbc01840-9ca6-42bb-a828-392175214f38 · inbound

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation cites this paper.

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-09T18:51:12.515735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:51:12.515735Z digest=sha256:5966a8f8d20086afa8befaf4ebc309c8ed56c67db6b705ef007d8e2b81fb5a7f

Observation be437bdc-f106-4352-adac-b851ecbde8da · inbound

Time Series Forecasting Through the Lens of Dynamics cites this paper.

Time Series Forecasting Through the Lens of Dynamics Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:32:01.305235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T03:31:48.513830Z digest=sha256:7d7451c683e2cf94187ec75670f6b48d2efeafd0a1302139678af96d42c33805