Pith. sign in

Paper Citation Record · LEDGER

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

As of 12 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 12 inbound Pith citation observations for arXiv:2506.01115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01115 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:22.908318Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:14:28.463703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.880747Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8363ba87-329c-42c9-9ecc-0a23659bea68 · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.571092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.571092Z digest=sha256:8eadeb10d2e8df0a744daca4ab73d709bfed12192f4bc9079022b136b5eed7a8

Observation e9abca2d-5c2b-440f-9692-f09073b1fa79 · outbound

This paper cites The Curious Case of Benign Memorization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The Curious Case of Benign Memorization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.702593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.577524Z digest=sha256:a7b30adee501f7456bfd8caa3e2a5af5474ef0e8673316e053e2e9d759f5e69a

Observation 73175888-b0e0-4f1b-a063-df84c3ac5ce3 · outbound

This paper cites Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang.On exact computation with an infinitely wide neural net.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang.On exact computation with an infinitely wide neural net

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.158136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.583250Z digest=sha256:9d3e7bc9876be14867a81bf97c7f64268c9ff5726196c78cfd67d2ec92513b48

Observation e41ea57a-c370-4585-bf0c-5511b8b9b17a · outbound

This paper cites A closer look at memorization in deep networks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A closer look at memorization in deep networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.143410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.588325Z digest=sha256:251726a0fb812dcd5308dbc6de0424b1ea7dcfc51a11b1e8fbd3bccc792a0fe9

Observation 117151e8-5cca-4bda-b64b-6e54be0079b1 · outbound

This paper cites Scaling mlps: A tale of inductive bias.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Scaling mlps: A tale of inductive bias

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.123851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.593277Z digest=sha256:71e60d6dd92cec238b9bfb7f4567ba7ad9f5759b6f312f30bb6d2387cf4b2d58

Observation 879bb15d-bee3-410b-a93f-0189690761f1 · outbound

This paper cites Mechanistic interpretability for AI safety - a review.Transactions on Machine Learning Research, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic interpretability for AI safety - a review.Transactions on Machine Learning Research, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.599011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.599011Z digest=sha256:d71a28372a5e70a410f09a59ee7629272ad4c5034f6247937d1a158b167fd3c2

Observation 0ea2cbf4-4bb5-4b85-af8e-f660e06859a4 · outbound

This paper cites Birth of a transformer: A memory viewpoint.Advances in Neural Information Processing Systems, 36:1560–1588, 2023.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Birth of a transformer: A memory viewpoint.Advances in Neural Information Processing Systems, 36:1560–1588, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.096487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.605192Z digest=sha256:cc9c11eca7561b4d42f96ebf78db3826378f45e0bfc1686965b1d0a736d01bcc

Observation 662cb06e-280c-496c-b623-4dc346d22452 · outbound

This paper cites Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.631525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.611391Z digest=sha256:4e3e0cf853ff5206d0a8f233f48e362a005e4b2eaba2f8b8c0d3fead67cdebed

Observation 3191c054-6843-42be-a1a1-84be8dbd1f98 · outbound

This paper cites Transformers generalize differently from information stored in context vs in weights.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers generalize differently from information stored in context vs in weights

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.619625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.619625Z digest=sha256:010b3e40f9db6b01c3cd66191c533694cb0bd3154182432e678b010135ec053c

Observation 30c7a563-58a0-4209-ae7e-594a24231996 · outbound

This paper cites Distributional associations vs in-context reasoning: A study of feed-forward and attention layers.ICLR, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Distributional associations vs in-context reasoning: A study of feed-forward and attention layers.ICLR, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.082310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.626924Z digest=sha256:60cd6179d6ee844337448e5d19ee8b7d1c7494d9f510aa1124fc5ac0692d3f5f

Observation 68fa1850-9619-4b11-837e-0cf11690ca00 · outbound

This paper cites Knowledge localization: Mission not accomplished? enter query localization! InProceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge localization: Mission not accomplished? enter query localization! InProceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.068319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.633098Z digest=sha256:6cbd659e660df1f915ed9f09068157f3c0b766c6d77521b311d470ca98a0c22a

Observation 8ca9d54d-76bd-4019-af87-fba6bb10bc3e · outbound

This paper cites Summing up the facts: Additive mechanisms behind factual recall in llms, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Summing up the facts: Additive mechanisms behind factual recall in llms, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.054015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.639502Z digest=sha256:8ee29011963e32090fc2db0a86c7f8120975607d94832210aa76ae285dfa47f8

Observation 1e846ff2-6ed8-4deb-8034-b1fb20869d90 · outbound

This paper cites Induction heads as an essential mechanism for pattern matching in in-context learning.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Induction heads as an essential mechanism for pattern matching in in-context learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.040339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.647499Z digest=sha256:fc41250d88ca9dca4a05fdb334b34323be5345116b2115435abee1ccc8700220

Observation 83408a97-f0ab-447b-a5e6-c01f0c6a15c1 · outbound

This paper cites Knowledge neurons in pretrained transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge neurons in pretrained transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.655162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.655162Z digest=sha256:3286cce7bbe10bc2a2dcf4c37180b41f6555448bed078574861a24dfd31bf954

Observation 44cc6a85-ddda-418d-9e36-2645fb4adb39 · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.027077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.661558Z digest=sha256:89a8eae6e853f8157a9dc43b8e56d3be5a5332c743df7ed06e7a810281716ec1

Observation 28cc1dae-114a-43c7-a596-2b6d94072207 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.667162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.667162Z digest=sha256:53cfcb224919134cb2c87427501550f78da64be9b3101f65246acb22a728bfc0

Observation 1d511882-7077-4c50-b2ab-5aa70f144bfa · outbound

This paper cites Edelman, eran malach, and Surbhi Goel.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Edelman, eran malach, and Surbhi Goel

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.008710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.672227Z digest=sha256:0dc7f4a2c50655ac84d1f2a9aca50c5a3ea100cc302e72c8c16618f328f71981

Observation 6de929e4-298d-4f9f-bd46-dcf17694bbd6 · outbound

This paper cites A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.989975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.677623Z digest=sha256:509e1085e1f208b22f5082e14bf4e577e2b6eb893b6d97b09e8bd98730b112f9

Observation 58bf9a70-406d-4f25-8238-7eedfc9358da · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer Feed-Forward Layers Are Key-Value Memories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.682505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.682505Z digest=sha256:e084e90687789719906cbbac8f35f06d8169084afda88c0cfe25be93ba11b11e

Observation c47962c0-f560-41db-b6ee-fa5f74c6d609 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers are key-value memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.687732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.687732Z digest=sha256:6a1f5e620aa8b4c345969e7832f9b6119aecf94d71a2dc0c2e519ecb86ae0423

Observation c0f87075-542c-4a28-afc8-718b2ecba1d3 · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.692163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.692163Z digest=sha256:cf265f636f744c49d1c4893d4ec97126626fa5800248bb6b7debf063b973440f

Observation 2dc95dbb-4299-4414-a852-483b5d696015 · outbound

This paper cites Dissectingrecalloffactualassociations inauto-regressivelanguagemodels.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Dissectingrecalloffactualassociations inauto-regressivelanguagemodels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.696572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.696572Z digest=sha256:386ef83983483200ce489d35bbcfe3ab2bdbb94bba30cb8c2d3fdc379d51a806

Observation a3e0fa06-aee0-48fa-b920-79811df9a4f2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.701367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.701367Z digest=sha256:cb7b3d6c60c032cdb98bd1a8ce744e43f42a127ccc5f00fa8df6381361530049

Observation 9631f467-18cd-46a7-aa33-fb364dfc04cf · outbound

This paper cites Smith, and Roy Schwartz.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Smith, and Roy Schwartz

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.974388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.705870Z digest=sha256:d63ad3ae64df1f62047774213ea97835ba16a79a27b12f373e597972d8ea14c5

Observation 2849f7ac-531f-4d2e-9980-80ffde91facd · outbound

This paper cites Simplifying Transformer Blocks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Simplifying Transformer Blocks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.710780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.710780Z digest=sha256:fd63527522461242868c5d4b960bd6fd189a74e04b6b3721468d43de95d4ed40

Observation 20848180-274a-4d69-a981-68c6b2675f35 · outbound

This paper cites Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.716704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.716704Z digest=sha256:5312ac300a9e68870ad66b0304549e9c0a11d4cbdc3e76370bfaad8d25330cb2

Observation 547444a1-677d-4f56-a61f-815b88a3eef3 · outbound

This paper cites What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.722380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.722380Z digest=sha256:b12adfeff3da4f9252afaf195cec0656857eed4180161da5219da0c2085a4f3b

Observation 657d772e-82b4-4721-82b6-807cdc2c3692 · outbound

This paper cites Lexico: Extreme kv cache compression via sparse coding over universal dictionaries, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Lexico: Extreme kv cache compression via sparse coding over universal dictionaries, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.960220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.728573Z digest=sha256:aa4cc7a790fcebeb0fc9b1a1aa64d84d20e414727f024401aeb0d4a99ea41579

Observation 1487f04a-6d72-4d38-87fd-03cc49bd03a7 · outbound

This paper cites Deep Neural Networks as Gaussian Processes.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Neural Networks as Gaussian Processes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.735092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.735092Z digest=sha256:01a7794fa23c91a672753b8b0fbae6fd5d879d27a9555aea6220c8c398ea0ed4

Observation 712c06fc-030a-4711-a0f1-dbb80d30ae53 · outbound

This paper cites FNet: Mixing Tokens with Fourier Transforms.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer FNet: Mixing Tokens with Fourier Transforms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.740900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.740900Z digest=sha256:b2faae42e17dd1afee0413ba8e136888d7be3c4830134d3ead86ecdd257c285a

Observation 5c2b73b7-af45-463d-9a16-cae3f6618d03 · outbound

This paper cites The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795–10808, 2022.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795–10808, 2022

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.946050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.745978Z digest=sha256:d50d3b20671d14a940a28c220734ffba005a74f2ab38be615f99041546a8250c

Observation 59b06326-5916-47b7-9a17-d49c8b382d79 · outbound

This paper cites Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.750524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.750524Z digest=sha256:d4ff6276d82e0e61a67828504570af5a9a3efcd977632ffc0e7c2c54512f90ce

Observation f8d18586-9ccc-4f75-92cb-096e3db23870 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating and Editing Factual Associations in GPT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.755732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.755732Z digest=sha256:ba39c2e4fdc301f404872bfea4358906b884ce1808ca7de7fa78f719323eeff7

Observation a0aca01b-a1f7-4fd7-a5da-eabfdb3d31f5 · outbound

This paper cites Pointer Sentinel Mixture Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Pointer Sentinel Mixture Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.760917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.760917Z digest=sha256:f37c8ebde10124364b9192097fa6217065314a768833c39e1b2c928116cf1dc2

Observation 1565742c-8d44-47e6-a5d7-46ce1da0bc03 · outbound

This paper cites Language models implement simple word2vec-style vector arithmetic, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Language models implement simple word2vec-style vector arithmetic, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.927859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.765913Z digest=sha256:94e8416b838752accdbb256405f44dd9e91d7df356cabca0baaadc0fd03c9aec

Observation 80ff018a-334f-4a56-abb3-99e13bcdbd6a · outbound

This paper cites Universal approximation property of random neural networks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Universal approximation property of random neural networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.911077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.771350Z digest=sha256:aa3e36536568feb46dd3a2971a39b53cd0f850406c2fb37c557a05101e41d0e3

Observation 21a9c1d6-595c-4cb0-9df0-32949282d797 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.Advances in Neural Information Processing Systems, 35:27198–27211, 2022.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.Advances in Neural Information Processing Systems, 35:27198–27211, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.775855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.775855Z digest=sha256:6fa1c357ff53c92c31e38beaec8a50476a84abfdd1336604c01b811dc5b8a561

Observation 2d5c8a7b-2e2e-4611-8957-ff5225150e0b · outbound

This paper cites The shaped transformer: Attention models in the infinite depth-and-width limit.Advances in Neural Information Processing Systems, 36:54250–54281, 2023.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The shaped transformer: Attention models in the infinite depth-and-width limit.Advances in Neural Information Processing Systems, 36:54250–54281, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.880239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.780004Z digest=sha256:a90ac48e45ddbb848db5c002ea47e4b6072894208b5be847c00a046550454737

Observation dbacfc98-69da-4b05-ac89-594e5b700859 · outbound

This paper cites Investigating the Limitations of Transformers with Simple Arithmetic Tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Investigating the Limitations of Transformers with Simple Arithmetic Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.784025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.784025Z digest=sha256:c1a89e323d69a67d8284ea07284cda50d261bc903b9e42e08cf79b011ed71873

Observation 0211457a-d6c4-426d-b890-b1f416dfa5b7 · outbound

This paper cites In-context Learning and Induction Heads.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer In-context Learning and Induction Heads

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.788567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.788567Z digest=sha256:f99a8229156f7b651a3bd8cf633f464eb56a23c20a92eade405067b7cd8df0b3

Observation 6a25774c-81d1-49bd-bc77-47473673ab53 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.793335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.793335Z digest=sha256:48b92c6bb8929bde2376f21e23367ef824b49d066fe8d43717f7df8444419298

Observation 7dcd428a-9b5d-4e5f-987c-30bbdeab8320 · outbound

This paper cites Mechanistic Design and Scaling of Hybrid Architectures.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic Design and Scaling of Hybrid Architectures

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.798604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.798604Z digest=sha256:05c01a7485c9a40bae78e794a5a3e50608f2faee228c32c919e73f1091a32812

Observation 021dc086-0b6c-4db4-8065-c4b2caa7fd2e · outbound

This paper cites Exponential expressivity in deep neural networks through transient chaos.Advances in neural information processing systems, 29, 2016.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Exponential expressivity in deep neural networks through transient chaos.Advances in neural information processing systems, 29, 2016

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.804007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.804007Z digest=sha256:ee57be00baf7f340712208ba88083a4c297e7bf79f2da04daf3b9a6eaf16a09a

Observation 5569d08c-f9ac-48ce-88af-7ae4d248fd97 · outbound

This paper cites Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.808829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.808829Z digest=sha256:99de8bf833198240c4b8f789da9397d5e292e0d39b479ff5d341a4c1ab0cf6b6

Observation dd623bde-0640-418a-83e6-52d2867d6da9 · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers, parallel computation, and logarithmic depth

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.814818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.814818Z digest=sha256:8c10fa3b12a651d3dc1f887ef64030b3002c6c6b85287df0bfcbdba2da432aff

Observation 11d2c10c-c4bd-4ca0-97d0-bf2ca5d19531 · outbound

This paper cites Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.855012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.819921Z digest=sha256:16897276df9b4d07dec435af414976abd6c74424c675e117b1db47c8c6e3a443

Observation 4aaf9a34-45f3-46e3-bf4c-199425e2134a · outbound

This paper cites Deep Information Propagation.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Information Propagation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.825593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.825593Z digest=sha256:0f9864854bb27d0d6922e254706eed2533d51f77aa37ebb671c2ac3b87af050e

Observation 3eaae48b-ba2b-486d-a579-8517c8271da9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.830990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.830990Z digest=sha256:0d85a8e79298c931761179bc248ff01935ee1e9d5e4d811c20857ec231f25c1c

Observation a6875d19-b787-4d3f-a7f3-4611c9930bcd · outbound

This paper cites Synthesizer: Rethinking self-attention in transformer models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Synthesizer: Rethinking self-attention in transformer models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.827631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.836163Z digest=sha256:c3eff390d5f3adacd42f62f8f9b42d415e974edfc38854893151cb387ed03a21

Observation 63bf97f3-c7f4-429d-99be-880d433f5eb0 · outbound

This paper cites Efficient Transformers: A Survey.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient Transformers: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.842653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.842653Z digest=sha256:608279f83c93ca626dd130fda8c8d289a2f478e874f5b49b75352026e73e4be2

Observation b34f8d58-6c3e-4c0d-99b8-b42f47b9359d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.848934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.848934Z digest=sha256:8d3e4f538466aad7dd9d16747bc69136fa0835e567522bdc25b3bef5b4245b8d

Observation 0d52af51-62e5-487e-b502-cf2f62f01f67 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.853190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.853190Z digest=sha256:552bd06eccc227220538be2c00e87776d3ec5fee2aa754ec8e032c171dda529a

Observation dc020708-d608-41be-92e3-0f469d4e5e57 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.858670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.858670Z digest=sha256:939e837c32d772827b5e655bad65439e69edd4f8f5a9f6e433ddcc4d91012ae6

Observation d3defb31-0d36-4736-8d53-77758ef0d9be · outbound

This paper cites Efficient streaming language models with attention sinks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient streaming language models with attention sinks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.862990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.862990Z digest=sha256:0fb17909cb90b7a27e3176059d2d86663ee0e83150874cb5e23174a29f1454e4

Observation 096338b0-aec6-4fa4-8739-8d2059421669 · outbound

This paper cites On layer normalization in the transformer architecture.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer On layer normalization in the transformer architecture

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.867299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.867299Z digest=sha256:fb566ac10ddf3071f355e95e52b38c85a4a742e878b40480f458c787cbdb7d7e

Observation e41da55c-0d29-4096-a942-585c53b48eb9 · outbound

This paper cites Mean field residual networks: On the edge of chaos.Advances in neural information processing systems, 30, 2017.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mean field residual networks: On the edge of chaos.Advances in neural information processing systems, 30, 2017

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.782386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.872566Z digest=sha256:5e41ad200fb37a22ce408a252d717953e64fe2214a9f570011a6ef00219e760c

Observation 62a6dff9-fa4b-41a5-83b7-798521dfa5b2 · outbound

This paper cites Knowledge Circuits in Pretrained Transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge Circuits in Pretrained Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.877534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.877534Z digest=sha256:327b7c1e4db3e6f9366231117615d5f9ceb84d114aa913e1f49ce5f02a60e815

Observation 04f9a6a8-d1a4-4ed3-a4b2-cb39ecf1b7a0 · outbound

This paper cites Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.768403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.882846Z digest=sha256:8c7e1c5e3deecad9a8a581d0cb1874b7317cb871e15b2b4d03425b737c0f8095

Observation 56cf4fe1-77de-43a9-a710-6a90ec70ae58 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.753867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.887654Z digest=sha256:f13b561dea05a16d6c3194e65bc992bd6c197aed8658dfd7038ca7da0f58329b

Observation 82e33590-6882-4b6a-9e7f-16ecdf82cd1a · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Understanding deep learning requires rethinking generalization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.891916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.891916Z digest=sha256:61030bf41e5e34a2cd9a1250af09f8f25ec5bc528dc7840fab96ef925a8647d4

Observation f6072648-2491-47d1-9f03-c2373e9c0d1c · outbound

This paper cites Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.011772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.896639Z digest=sha256:00351455234591d08425957195e96902a0dc81c5b17bb343d244aed84477d067

Observation 294b2fae-c881-45c7-a025-2371793ba2e9 · outbound

This paper cites Character-level Convolutional Networks for Text Classification.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Character-level Convolutional Networks for Text Classification

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.901611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.901611Z digest=sha256:47545ede1243e63b8eac91967a6897cc45a294a5857e6c1d65fdeb0161b8b39c

Observation 8b1f9f34-206f-4e51-ae70-d8265d664326 · outbound

This paper cites Algorithmic capabilities of random transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Algorithmic capabilities of random transformers

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.740198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:57:22.908318Z digest=sha256:841f33251c31f6701239071e978940903fa08714d718b04c83720af7d76506a4

Pith citing papers

Observation 1a7b7cd7-c4d6-4a78-9f4e-10c11430db8b · inbound

Provable Knowledge Acquisition and Extraction in One-Layer Transformers cites this paper.

Provable Knowledge Acquisition and Extraction in One-Layer Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:10:44.852070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T23:05:56.687644Z digest=sha256:b23b1df69993902f102738eddcf7f35f3a31fa8df4bfb4a605344f4dd39186e5

Observation 25956d3d-2ece-4dbc-bcf9-d6d2a8a4b5f6 · inbound

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers cites this paper.

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:15:07.096284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:15:07.096284Z digest=sha256:e191dbefecaccf5f198c88a8eade6dbb3eef601dccd34748907c0e28d46fbac8

Observation b9087432-cd4e-42a1-8b5b-3fa872b1f2c0 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:19.053632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T22:19:25.483640Z digest=sha256:f64c3546f52a23313a2b31392ddb4ce0a6d89189786b43da0ca2e04b78354908

Observation 5cd2c773-03ed-4a2b-a62c-a709c6b164d2 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:54:51.174733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T11:54:29.436149Z digest=sha256:3736e4a06337e7ee30bb52842f47b832b5ca927775241b5cfea52e0c9fc08dcf

Observation 2ef29a84-3c61-4c99-a577-712726e622bb · inbound

Procedural Pretraining: Warming Up Language Models with Abstract Data cites this paper.

Procedural Pretraining: Warming Up Language Models with Abstract Data Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:55:10.047706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:55:10.047706Z digest=sha256:53d177aaf594cae4ce2d118f127c4a87f461cb72c4b2be57aeb056977910537c

Observation 5c3d438d-9afd-463d-93e2-61c7f1e8720e · inbound

Geometry-Calibrated Conformal Abstention for Language Models cites this paper.

Geometry-Calibrated Conformal Abstention for Language Models Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:28.657564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T05:56:03.031192Z digest=sha256:d963585a1569be2806dda3a2004ba2046ceac22c05ea9533050a3ac3043079bb

Observation 3103343d-d627-4c1d-b9a7-94dba112716b · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.807728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T11:47:41.717389Z digest=sha256:d0ca62bf0b578977e2df3cfa54e5b5e9fdfd712f1edada7842310b90dfef2c5a

Observation b314efdc-8adb-44b0-8ae0-3031b2ecdab7 · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:15:11.538789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T07:15:02.755541Z digest=sha256:4287d47cf1dde2eb090bde78eadd162c1b96636db69085f551c6747943d5539f

Observation 20037440-1bb0-4963-9551-f74d7b62b213 · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T11:17:06.440358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:17:06.440358Z digest=sha256:193330522b3628df6c984765717889c52e5abbaefd7a25ebb492999efb581fcd

Observation 028dea82-b222-470e-bee6-0ef6955c05fe · inbound

Fixed Universal Transformers cites this paper.

Fixed Universal Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.217756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:24:18.984754Z digest=sha256:c2bf69a88a27a61fe685bd66174447b4b43c6df408e7a22470e8d5dc60d0b157

Observation 8dbf2c27-2496-441f-bdef-32e954d60dcf · inbound

Activation-Based Active Learning for In-Context Learning: Challenges and Insights cites this paper.

Activation-Based Active Learning for In-Context Learning: Challenges and Insights Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.882218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T06:27:29.827899Z digest=sha256:23f95acd48769484a5a4c3389825ec4abccb4a2437121ba24c8dbadbe700741b

Observation 69de66da-b7a2-4dcd-93e9-190b1aee2527 · inbound

Training-Free Universal Approximation by Prompting Random Transformers cites this paper.

Training-Free Universal Approximation by Prompting Random Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:28.463703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:14:28.463703Z digest=sha256:8b9a091564667def4773ce0e1249b89dd4e5ed307e95ed009e174b18641327bf