REVIEW 3 major objections 6 minor 86 references
Converting Transformers into DGNNs Form
T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A neural layer that replaces self-attention with a learned digraph Fourier convolution beats all fifteen Transformer baselines on Long-Range Arena, averaging 75.94% accuracy.
desk verdict The LRA gain is unverified and likely an artifact of copied baselines; the Givens-rotation + kernel polynomial method is a real idea worth testing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is Synvolution, a learnable unitary digraph convolution defined as $\mathrm{Synv}(XW_V) = \Phi^{-1}\bigl(\exp(i\Lambda) \odot (\Phi XW_V)\bigr)$, where $\Phi$ is a synthetic unitary eigenvector matrix and $\exp(i\Lambda)$ is a diagonal matrix of synthetic eigenvalues; together they act as the digraph shift operator and frequency response of a digraph Fourier transform. The matrix $\Phi$ is built by the order-1 DHHP parametrization, a unitary diagonal matrix times one lower-unitary Hessenberg factor times one upper-unitary Hessenberg factor assembled from Givens rotations, so the transform runs in $O(N \log N)$ time via parallel scan rather than $O(N^2)$ matrix multiplication. Kernelution replaces the raw eigenvalue multiplier with a kernel-polynomial filter $p_{\mathrm{KP}}(\Lambda)$ whose learnable Chebyshev coefficients are penalized by the kernel polynomial loss, which simulates a data-dependent, Gibbs-damped kernel. A Gated Feed-Forward Network maps the complex-valued output back to real values, and PostScaleNorm stabilizes training.
What would settle it
Re-run all fifteen Long-Range Arena baselines under Converter's own protocol, the same Bayesian tuning budget, seeds, tokenization, and normalization, and check whether the 75.94% average and the fourteen-point gap over Luna-256 survive. Separately, measure the reconstruction error of 1-DHHP against random unitary and DFT matrices of growing size to test whether the Givens-Hessenberg product actually covers the unitary group as Assumption 1 claims.
Extended reading notes
Core claim
The paper's central claim is that a Transformer whose self-attention is swapped for a synthetic unitary digraph convolution outperforms the vanilla Transformer and fourteen efficient-Transformer variants on every one of the five Long-Range Arena tasks, with scores of 60.38% on ListOps, 86.44% on Text, 83.41% on Retrieval, 61.02% on Image, and 88.43% on Pathfinder, a 75.94% average that exceeds the second-best model, Luna-256 at 61.95%, by roughly fourteen points. The same Converter reports the best accuracy on arXiv-document classification at 16K and 32K tokens (81.77% and 82.34%) and on DNA taxonomy classification (84.59% on the Bos/Sus task and 59.49% on the Mus/Rattus task). The paper reads these results as evidence that the softmax similarity bottleneck is not necessary: a full-rank, dense, learnable mixing operator built from Givens rotations and spectral filtering can carry long-range dependency modeling by itself. Its ablations add that the GRU-based relative position embedding and the kernel polynomial loss each contribute substantially to the final accuracy.
Load-bearing premise
The headline performance claim depends on the fifteen Long-Range Arena baselines being tuned and evaluated under the same protocol as Converter, which the paper does not document (their numbers match published results), and the theoretical claim depends on the unproven Assumption 1 that at most $\lceil N/4 \rceil$ orders of the L-DHHP rotation product can represent any dense unitary matrix.
Editorial extensions
If this is right
- Self-attention is not required for Transformer-level performance: a linearithmic, full-rank, dense spectral mixing layer suffices on sequences up to 32K tokens.
- The kernel polynomial method can stand in for the multi-head operation, with the new kernel polynomial loss supplying a principled, order-increasing penalty that mimics adaptive Gibbs damping.
- Because the mixing matrix is unitary and full-rank by construction, Converter's layer avoids the low-rank bottleneck that limits softmax and kernel attention as depth grows.
- The same spectral layer transfers across raw text, flattened pixels, long documents, and DNA sequences, indicating the construction is domain-agnostic rather than tuned to one data type.
- A single-head spectral layer with linearithmic cost is enough to beat models built around sparse, low-rank, or kernelized attention on the tested benchmarks.
Reading between the lines
- If the Table 1 baselines were re-run under Converter's exact tuning protocol, the fourteen-point Long-Range Arena margin could shrink; the residual gap on ListOps and Image would then be the honest measure of the mechanism's advantage.
- The 1-DHHP fast transform is a general structured unitary layer, so it could slot into other quadratic-complexity positions such as cross-attention decoders or state-space mixers, which the paper names as future work for cross-attention.
- The kernel polynomial loss ties regularization strength to filter order; letting the order $K$ itself be scheduled or learned during training would be a natural extension from coarse to fine spectral resolution.
- Because the unitary constraint keeps the spectrum on the unit circle, deeper Converter stacks are plausibly more stable than unconstrained attention; a controlled depth-scaling study would test whether the spectral construction delays rank collapse.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Converter, a Transformer variant in which the self-attention module is replaced by a synthetic unitary digraph convolution called Synvolution, built on the order-L DHHP parametrization of unitary matrices and a learned diagonal of eigenvalues. The authors also introduce Kernelution, which applies the kernel polynomial method to the spectral filter, a kernel polynomial loss, and a gated feed-forward network for complex-valued activations. The paper claims linearithmic time complexity, full-rank dense mixing, and state-of-the-art accuracy on the Long-Range Arena benchmark, long document classification, and DNA taxonomy classification, with an average LRA accuracy of 75.94% versus 61.95% for the best baseline.
Significance. If the empirical results are reproducible under a matched protocol, this is a significant contribution: a simple, full-rank, dense attention alternative with linearithmic complexity that outperforms a wide range of Transformer variants on long-sequence benchmarks would be of broad interest. The theoretical framing via digraph signal processing is novel, and the L-DHHP construction with a fast parallel-scan implementation is a concrete algorithmic contribution. The paper also includes useful ablations showing the contribution of RPE and the kernel polynomial loss. However, the main significance currently rests on an empirical comparison whose fairness is not documented, and the central theoretical proposition is contingent on an unproven assumption.
major comments (3)
- [§5.1 and Appendix F.1] The headline claim that Converter surpasses all 14 baselines on all five LRA tasks by a 14-point average margin depends entirely on Table 1. Appendix F.1 states that 'all four datasets and seven baseline models' were tuned with Bayesian optimization under a 16 GB memory constraint, but Table 1 contains five LRA tasks and fifteen models, and Tables 7 and 8 provide no LRA baseline hyperparameters. The reported baseline numbers closely match those in Tay et al. (2021b), which suggests they were transcribed from the literature rather than re-run under the paper's protocol. As presented, the comparison is not validated: the tuning budget, seeds, preprocessing, and evaluation procedure for the baselines are unknown. Please re-run all baselines under the same protocol, report their hyperparameters, and provide the provenance of each value in Table 1, or substantially soften the superiority claim.
- [Appendix E, proof of Proposition 2] Proposition 2 is not derived but merely restated: the proof says that since DFT, DWHT, DCT, and DST are unitary matrices, Assumption 1 implies they can be represented by L-DHHP. Because Assumption 1 already asserts that L-DHHP can construct every dense unitary matrix, the proposition carries no independent content. The assumption itself is nontrivial and unproven, and the paper gives no constructive procedure or numerical evidence that ⌈N/4⌉ orders suffice. Please either prove Assumption 1 or a weaker version that covers the specific transforms, or explicitly mark the theoretical claims as conditional.
- [§5 and Appendix F.1] No seed information or error bars are reported anywhere in the experiments, and no code repository link is provided. LRA results are known to be sensitive to initialization and tuning, so single-run accuracy values without seeds are insufficient to support a 14-point margin. Please report the number of runs, the variance, and release the code with a clear reproducibility script; this is a load-bearing issue for the empirical claims.
minor comments (6)
- [§4.2] The text uses both 'Kernolution' and 'Kernelution' (and later 'Synolution'); please standardize the spelling.
- [§5.4] The ablation text says 'Converter achieves the highest performance when using PRE, followed by APE and SPE', but the table shows no 'PRE' row; this should be 'RPE', the recurrent position embedding used elsewhere.
- [§5.1] The text describes LRA as containing 'five multi-class classification tasks', but Text and Retrieval are binary classification tasks; please correct the wording.
- [Appendix F.1] The sentence 'For all four datasets and seven baseline models' contradicts the actual experimental scope; the appendix should accurately state the number of datasets and baselines.
- [References] The reference list contains a typo in 'Elena V oita' (a stray space) and should be corrected.
- [Equation (7)] The product notation in Equation (7) is difficult to read and appears malformed in the preprint; please use explicit index bounds or a clearer notation.
Circularity Check
The only reduce-by-construction step is Proposition 2, which is a direct instantiation of unproven Assumption 1; the empirical comparison is external and not circular.
-
other
[Section 4.1 (Assumption 1 / Proposition 2) and Appendix E (Proof of Proposition 2)]
"Since our method is based on the Givens rotation method, we make Assumption 1. Under this assumption, we can establish the following propositions. Assumption 1. For constructing an arbitrary N × N dense unitary matrix, at most ⌈ N/4 ⌉ orders are sufficient for L-DHHP. Proposition 2. L-DHHP captures the discrete unitary transforms, including discrete Fourier transform (DFT), the discrete Walsh–Hadamard transform (DWHT), the discrete cosine transform (DCT), the discrete sine transform (DST), and their inverses exactly."
The proof in Appendix E reduces Proposition 2 to Assumption 1: 'By Assumption 1, any N × N unitary matrix can be exactly constructed using L-DHHP... Therefore, as specific cases of unitary matrices, DFT, DWHT, DCT, and DST can all be exactly represented by L-DHHP.' Since Assumption 1 already asserts that every dense unitary matrix is representable by L-DHHP, the proposition is a one-line instantiation of the assumption together with the definitional fact that these transforms are unitary. The claimed expressive-completeness result is assumed rather than independently derived. The paper is transparent that this is conditional, so this is a minor tautological step rather than a hidden circularity, but the proposition adds no evidence beyond the assumption.
full rationale
The central performance claims are tested against external benchmarks (LRA, long-document, and DNA taxonomy), so they are not circular: no baseline accuracy or target value is used as a fitted input, and no parameter fitted to one subset is relabeled as a prediction on a closely related subset. There are no load-bearing self-citations by the authors. The only step that reduces to its own input by construction is Proposition 2, which is a direct corollary of unproven Assumption 1; however, the empirical results do not depend on this proposition, and the paper explicitly labels the premise as an assumption. Protocol concerns about LRA baseline tuning (Appendix F.1) are a correctness and validity risk, not a circularity. Overall score 2 reflects this single auxiliary tautological theoretical step while the main empirical claim is independently benchmarked.
Assumptions & free parameters
free parameters (3)
- K, maximum Chebyshev order =
2 (all tasks)
- eta, KPL weight =
0.001 for ListOps, Text, Retrieval, Pathfinder; 0.01 for Image; 0.1 for LongDoc and Ensembl
- DHHP order L =
1
assumptions (4)
- ad hoc to paper Assumption 1: at most ceil(N/4) orders of L-DHHP suffice for any N x N dense unitary matrix
- standard math Chebyshev polynomial interpolation theorems for differentiable and analytic functions
- standard math Kernel polynomial method with Gibbs damping factors
- standard math LQ decomposition of a square matrix into a lower triangular matrix and Givens rotations
Cite this review
Pith. "Pith review of Converting Transformers into DGNNs Form." pith.science (2026). https://pith.science/paper/6DIS53JX
@misc{pith2026250200585,
author = {Pith},
title = {Pith review of: Converting Transformers into DGNNs Form},
year = {2026},
howpublished = {\url{https://pith.science/paper/6DIS53JX}},
note = {Machine review of arXiv:2502.00585}
}
read the original abstract
Recent advances in deep learning have established Transformer architectures as the predominant modeling paradigm. Central to the success of Transformers is the self-attention mechanism, which scores the similarity between query and key matrices to modulate a value matrix. This operation bears striking similarities to digraph convolution, prompting an investigation into whether digraph convolution could serve as an alternative to self-attention. In this study, we formalize this concept by introducing a synthetic unitary digraph convolution based on the digraph Fourier transform. The resulting model, which we term Converter, effectively converts a Transformer into a Directed Graph Neural Network (DGNN) form. We have tested Converter on Long-Range Arena benchmark, long document classification, and DNA sequence-based taxonomy classification. Our experimental results demonstrate that Converter achieves superior performance while maintaining computational efficiency and architectural simplicity, which establishes it as a lightweight yet powerful Transformer variant.
Figures
Reference graph
Works this paper leans on
-
[1]
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention Is All You Need . In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30, pages 5998--6008. Curran Associates, I...
2017
-
[2]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, V...
2019
-
[3]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale . In International Conference on Learning Representations, 2021
2021
-
[4]
Transforming the Language of Life: Transformer Neural Networks for Protein Prediction Tasks
Ananthan Nambiar, Maeve Heflin, Simon Liu, Sergei Maslov, Mark Hopkins, and Anna Ritz. Transforming the Language of Life: Transformer Neural Networks for Protein Prediction Tasks . In Proceedings of the 11th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics. Association for Computing Machinery, 2020
2020
-
[5]
Transformer Dissection: An Unified Understanding for Transformer ' s Attention via the Lens of Kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. Transformer Dissection: An Unified Understanding for Transformer ' s Attention via the Lens of Kernel . In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing...
2019
-
[6]
Rethinking Attention with Performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. Rethinking Attention with Performers . In International Conference on Learning Representations, 2021
2021
-
[7]
cosFormer: Rethinking Softmax In Attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong. cosFormer: Rethinking Softmax In Attention . In International Conference on Learning Representations, 2022
work page 2022
-
[8]
Attention is not all you need: pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: pure attention loses rank doubly exponentially with depth. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 2793--2803. PMLR, 07 2021
work page 2021
Show all 86 references
-
[9]
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. Breaking the softmax bottleneck: A high-rank RNN language model. In International Conference on Learning Representations, 2018
2018
-
[10]
softmax is not enough (for sharp out-of-distribution)
Petar Veli c kovi\' c , Christos Perivolaropoulos, Federico Barbero, and Razvan Pascanu. softmax is not enough (for sharp out-of-distribution). arXiv preprint arXiv: 2410.01104, 2024
2024 arXiv
-
[11]
Synthesizer: Rethinking Self-Attention for Transformer Models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. Synthesizer: Rethinking Self-Attention for Transformer Models . In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings ...
2021
-
[12]
FNet: Mixing Tokens with Fourier Transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. FNet: Mixing Tokens with Fourier Transforms . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4296--43...
2022
-
[13]
Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-Attention
Tong Yu, Ruslan Khalitov, Lei Cheng, and Zhirong Yang. Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-Attention . In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 681--690, 2022 a
2022
-
[14]
Big Bird: Transformers for Longer Sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big Bird: Transformers for Longer Sequences . In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Li...
2020
-
[15]
Silver and H
R.N. Silver and H. R \" o der. Densities of States of Mega-Dimensional Hamiltonian Matrices . International Journal of Modern Physics C, 05 0 (04): 0 735--753, 1994
1994
-
[16]
Calculating the density of states and optical-absorption spectra of large quantum systems by the plane-wave moments method
Lin-Wang Wang. Calculating the density of states and optical-absorption spectra of large quantum systems by the plane-wave moments method. Physical Review B, 49: 0 10154--10158, 04 1994
1994
-
[17]
Dielectric Constants of Silicon Quantum Dots
Lin-Wang Wang and Alex Zunger. Dielectric Constants of Silicon Quantum Dots . Physical Review Letters, 73: 0 1039--1042, 08 1994
1994
-
[18]
Kouri, and David K
Amrendra Vijay, Donald J. Kouri, and David K. Hoffman. Scattering and Bound States: A Lorentzian Function-Based Spectral Filter Approach . The Journal of Physical Chemistry A, 108 0 (41): 0 8987--9003, 10 2004
2004
-
[19]
The kernel polynomial method
Alexander Wei e, Gerhard Wellein, Andreas Alvermann, and Holger Fehske. The kernel polynomial method. Reviews of Modern Physics, 78: 0 275--306, 05 2006
2006
-
[20]
Chebyshev Expansion Techniques , pages 545--577
Alexander Wei e and Holger Fehske. Chebyshev Expansion Techniques , pages 545--577. Springer Berlin Heidelberg, 2008
2008
-
[21]
Long Range Arena: A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long Range Arena: A Benchmark for Efficient Transformers . In International Conference on Learning Representations, 2021 b
2021
-
[22]
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating Long Sequences with Sparse Transformers . arXiv preprint arXiv: 1904.10509, 2019
1904 arXiv
-
[23]
Reformer: The Efficient Transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The Efficient Transformer . In International Conference on Learning Representations, 2020
2020
-
[24]
Scatterbrain: Unifying sparse and low-rank attention
Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher R\' e . Scatterbrain: Unifying sparse and low-rank attention. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, vol...
2021
-
[25]
MetaFormer Baselines for Vision
Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, and Xinchao Wang. MetaFormer Baselines for Vision . arXiv preprint arXiv: 2210.13452, 2022 b
2022 arXiv
-
[26]
MetaFormer Is Actually What You Need for Vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. MetaFormer Is Actually What You Need for Vision . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10819--10829, 06 2022 c
2022
-
[27]
Are Sixteen Heads Really Better than One? In H
Paul Michel, Omer Levy, and Graham Neubig. Are Sixteen Heads Really Better than One? In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32, pages 14014--14024. Curran Asso...
2019
-
[28]
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797...
2019
-
[29]
Multi-Head Attention: Collaborate Instead of Concatenate
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. Multi-Head Attention: Collaborate Instead of Concatenate . arXiv preprint arXiv: 2006.16362, 2020
2006 arXiv
-
[30]
Low-Rank Bottleneck in Multi-head Attention Models
Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. Low-Rank Bottleneck in Multi-head Attention Models . In Hal Daum \' e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Procee...
2020
-
[31]
Graph filters for signal processing and machine learning on graphs
Elvin Isufi, Fernando Gama, David I Shuman, and Santiago Segarra. Graph filters for signal processing and machine learning on graphs. IEEE Transactions on Signal Processing, pages 1--32, 2024
2024
-
[32]
Aliaksei Sandryhaila and Jos \' e M. F. Moura. Discrete Signal Processing on Graphs . IEEE Transactions on Signal Processing, 61 0 (7): 0 1644--1656, 2013 a
2013
-
[33]
Aliaksei Sandryhaila and Jos \' e M. F. Moura. Discrete signal processing on graphs: Graph fourier transform . In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 6167--6170, 2013 b
2013
-
[34]
Aliaksei Sandryhaila and Jos \' e M. F. Moura. Discrete Signal Processing on Graphs: Frequency Analysis . IEEE Transactions on Signal Processing, 62 0 (12): 0 3042--3054, 2014
2014
-
[35]
Rahul Singh, Abhishek Chakraborty, and B. S. Manoj. Graph Fourier transform based on directed Laplacian . In 2016 International Conference on Signal Processing and Communications (SPCOM), pages 1--5, 2016
2016
-
[36]
Laplacians and the cheeger inequality for directed graphs
Fan Chung. Laplacians and the cheeger inequality for directed graphs. Annals of Combinatorics, 9 0 (1): 0 1--19, 04 2005
2005
-
[37]
Ala \' i z, and Johan A
Micha \" e l Fanuel, Carlos M. Ala \' i z, and Johan A. K. Suykens. Magnetic eigenmaps for community detection in directed networks. Physical Review E, 95: 0 022302, 02 2017
2017
-
[38]
Ala \' i z, \' A ngela Fern \' a ndez, and Johan A.K
Micha \" e l Fanuel, Carlos M. Ala \' i z, \' A ngela Fern \' a ndez, and Johan A.K. Suykens. Magnetic Eigenmaps for the visualization of directed networks . Applied and Computational Harmonic Analysis, 44 0 (1): 0 189--199, 2018
2018
-
[39]
Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform
Nir Ailon and Bernard Chazelle. Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform . In Proceedings of the Thirty-Eighth Annual ACM Symposium on Theory of Computing, STOC '06, pages 557--563. Association for Computing Machinery, 2006
2006
-
[40]
A sparse Johnson: Lindenstrauss transform
Anirban Dasgupta, Ravi Kumar, and Tam\' a s Sarlos. A sparse Johnson: Lindenstrauss transform . In Proceedings of the Forty-Second ACM Symposium on Theory of Computing, STOC '10, pages 341--350. Association for Computing Machinery, 2010
2010
-
[41]
Fastfood — Approximating Kernel Expansions in Loglinear Time
Quoc Le, Tamas Sarlos, and Alexander Smola. Fastfood — Approximating Kernel Expansions in Loglinear Time . In Sanjoy Dasgupta and David McAllester, editors, Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Resea...
2013
-
[42]
Orthogonal Random Features
Felix Xinnan X Yu, Ananda Theertha Suresh, Krzysztof M Choromanski, Daniel N Holtmann-Rice, and Sanjiv Kumar. Orthogonal Random Features . In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 29, pages...
1975
-
[43]
Deep Fried Convnets
Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, and Ziyu Wang. Deep Fried Convnets . In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 12 2015
2015
-
[44]
ACDC: A Structured Efficient Linear Layer
Marcin Moczulski, Misha Denil, Jeremy Appleyard, and Nando de Freitas. ACDC: A Structured Efficient Linear Layer . In International Conference on Learning Representations, 2016
2016
-
[45]
Hammond, Pierre Vandergheynst, and R \' e mi Gribonval
David K. Hammond, Pierre Vandergheynst, and R \' e mi Gribonval. Wavelets on graphs via spectral graph theory. Applied and Computational Harmonic Analysis, 30 0 (2): 0 129--150, 2011
2011
-
[46]
Implicit Neural Representations with Periodic Activation Functions
Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit Neural Representations with Periodic Activation Functions . In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing S...
2020
-
[47]
Computation of Plain Unitary Rotations Transforming a General Matrix to Triangular Form
Wallace Givens. Computation of Plain Unitary Rotations Transforming a General Matrix to Triangular Form . Journal of the Society for Industrial and Applied Mathematics, 6 0 (1): 0 26--50, 1958
1958
-
[48]
Learning Latent Permutations with Gumbel-Sinkhorn Networks
Gonzalo Mena, David Belanger, Scott Linderman, and Jasper Snoek. Learning Latent Permutations with Gumbel-Sinkhorn Networks . In International Conference on Learning Representations, 2018
2018
-
[49]
Monarch: Expressive Structured Matrices for Efficient and Accurate Training
Tri Dao, Beidi Chen, Nimit S Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher R \' e . Monarch: Expressive Structured Matrices for Efficient and Accurate Training . In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csa...
2022
-
[50]
Sparse factorization of square matrices with application to neural attention modeling
Ruslan Khalitov, Tong Yu, Lei Cheng, and Zhirong Yang. Sparse factorization of square matrices with application to neural attention modeling. Neural Networks, 152: 0 160--168, 2022
2022
-
[51]
Fast Training of Convolutional Networks through FFTs
Micha \"e l Mathieu, Mikael Henaff, and Yann LeCun. Fast Training of Convolutional Networks through FFTs . International Conference on Learning Representations, 2013
2013
-
[52]
Spectral Graph Theory
Fan Chung. Spectral Graph Theory . American Mathematical Society, 1997
1997
-
[53]
Trefethen
Lloyd N. Trefethen. Approximation Theory and Approximation Practice, Extended Edition . Society for Industrial and Applied Mathematics, 2019
2019
-
[54]
Edwin Hewitt and Robert E. Hewitt. The Gibbs-Wilbraham phenomenon: An episode in fourier analysis . Archive for History of Exact Sciences, 21 0 (2): 0 129--160, 06 1979
1979
-
[55]
Wong, and Lidia S
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao. Learning Deep Transformer Models for Machine Translation . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1810--1822. Association for ...
2019
-
[56]
Nguyen and Julian Salazar
Toan Q. Nguyen and Julian Salazar. Transformers without Tears: Improving the Normalization of Self-Attention . In Proceedings of the 16th International Conference on Spoken Language Translation. Association for Computational Linguistics, 11 2019
2019
-
[57]
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu ...
2019
-
[58]
On the Relation between Position Information and Sentence Length in Neural Machine Translation
Masato Neishi and Naoki Yoshinaga. On the Relation between Position Information and Sentence Length in Neural Machine Translation . In Mohit Bansal and Aline Villavicencio, editors, Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 32...
2019
-
[59]
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization . In International Conference on Learning Representations, 2019
2019
-
[60]
Peters, and Arman Cohan
Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The Long-Document Transformer . arXiv preprint arXiv: 2004.05150, 2020
2004 arXiv
-
[61]
Li, Madian Khabsa, Han Fang, and Hao Ma
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-Attention with Linear Complexity . arXiv preprint arXiv: 2006.04768, 2020
2006 arXiv
-
[62]
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran c ois Fleuret. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention . In Hal Daum \' e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, vol...
2020
-
[63]
Sparse Sinkhorn Attention
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. Sparse Sinkhorn Attention . In Hal Daum \' e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 9438-...
2020
-
[64]
o mformer: A Nystr \
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. Nystr \" o mformer: A Nystr \" o m-based Algorithm for Approximating Self-Attention . Proceedings of the AAAI Conference on Artificial Intelligence, 35 0 (16): 0 14138--14148...
2021
-
[65]
Luna: Linear Unified Nested Attention
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer. Luna: Linear Unified Nested Attention . In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021
2021
-
[66]
ListOps: A Diagnostic Dataset for Latent Tree Learning
Nikita Nangia and Samuel Bowman. ListOps: A Diagnostic Dataset for Latent Tree Learning . In Silvio Ricardo Cordeiro, Shereen Oraby, Umashanthi Pavalanathan, and Kyeongmin Rim, editors, Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Co...
2018
-
[67]
Maas, Raymond E
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning Word Vectors for Sentiment Analysis . In Dekang Lin, Yuji Matsumoto, and Rada Mihalcea, editors, Proceedings of the 49th Annual Meeting of the Association for Computational...
2011
-
[68]
Radev, Pradeep Muthukrishnan, and Vahed Qazvinian
Dragomir R. Radev, Pradeep Muthukrishnan, and Vahed Qazvinian. The ACL Anthology Network Corpus . In Min-Yen Kan and Simone Teufel, editors, Proceedings of the 2009 Workshop on Text and Citation Analysis for Scholarly Digital Libraries ( NLPIR 4 DL ) , pages 54--61. Associatio...
2009
-
[69]
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky. Learning Multiple Layers of Features from Tiny Images . Technical report, 2009
2009
-
[70]
Learning long-range spatial dependencies with horizontal gated recurrent units
Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre. Learning long-range spatial dependencies with horizontal gated recurrent units. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural I...
2018
-
[71]
Disentangling neural mechanisms for perceptual grouping
Junkyung Kim, Drew Linsley, Kalpit Thakkar, and Thomas Serre. Disentangling neural mechanisms for perceptual grouping. In International Conference on Learning Representations, 2020
2020
-
[72]
Parallel and serial grouping of image elements in visual perception
Roos Houtkamp and Pieter R Roelfsema. Parallel and serial grouping of image elements in visual perception. J. Exp. Psychol. Hum. Percept. Perform., 36 0 (6): 0 1443--1459, 12 2010
2010
-
[73]
Long length document classification by local convolutional feature aggregation
Liu Liu, Kaile Liu, Zhenghai Cong, Jiali Zhao, Yefei Ji, and Jun He. Long length document classification by local convolutional feature aggregation. Algorithms, 11 0 (8), 2018
2018
-
[74]
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. Convolutional Sequence to Sequence Learning . In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Le...
2017
-
[75]
Transformer Language Models without Positional Encodings Still Learn Positional Information
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy. Transformer Language Models without Positional Encodings Still Learn Positional Information . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 1382--1390. Association for Computational L...
2022
-
[76]
Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings
Ta-Chung Chi, Ting-Han Fan, Li-Wei Chen, Alexander Rudnicky, and Peter Ramadge. Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings . In Proceedings of the 61st Annual Meeting of the Association for Compu...
2023
-
[77]
The Impact of Positional Encoding on Length Generalization in Transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. The Impact of Positional Encoding on Length Generalization in Transformers . In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Info...
2023
-
[78]
Choose a Transformer: Fourier or Galerkin
Shuhao Cao. Choose a Transformer: Fourier or Galerkin . In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021
2021
-
[79]
Efficient Attention: Attention With Linear Complexities
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. Efficient Attention: Attention With Linear Complexities . In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 3531--3539, 01 2021
2021
-
[80]
Sparse Attention with Linear Units
Biao Zhang, Ivan Titov, and Rico Sennrich. Sparse Attention with Linear Units . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6507--6520. Association for Computational Linguistics, 11 2021
2021
-
[81]
SimA: Simple Softmax-Free Attention for Vision Transformers
Soroush Abbasi Koohpayegani and Hamed Pirsiavash. SimA: Simple Softmax-Free Attention for Vision Transformers . In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 2607--2617, 01 2024
2024
-
[82]
Skyformer: Remodel Self-Attention with Gaussian Kernel and Nystr \" o m Method
Yifan Chen, Qi Zeng, Heng Ji, and Yun Yang. Skyformer: Remodel Self-Attention with Gaussian Kernel and Nystr \" o m Method . In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pa...
2021
-
[83]
Scalable Parallel Programming with CUDA: Is CUDA the parallel programming model that application developers have been waiting for? Queue, 6 0 (2): 0 40--53, 03 2008
John Nickolls, Ian Buck, Michael Garland, and Kevin Skadron. Scalable Parallel Programming with CUDA: Is CUDA the parallel programming model that application developers have been waiting for? Queue, 6 0 (2): 0 40--53, 03 2008
2008
-
[84]
Untersuchungen \"u ber Fouriersche Reihen
Leopold Fej \'e r. Untersuchungen \"u ber Fouriersche Reihen . Mathematische Annalen, 58: 0 51--69, 1904
1904
-
[85]
Discourse on Fourier series
Cornelius Lanczos. Discourse on Fourier series. University mathematical monographs. Oliver and Boyd, 1966
1966
-
[86]
Veki \' c and S
M. Veki \' c and S. R. White. Smooth boundary conditions for quantum lattice systems. Physical Review Letters, 71: 0 4283--4286, 12 1993
1993
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.