Pith. sign in

Paper Citation Record · LEDGER

Two Heads Are Better than One: Simulating Large Transformers with Small Ones

As of 17 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12220 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:12:48.259338Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:05:58.230449Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b51397b-9819-4a56-bba0-0277ff4eb8cd · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Zoology: Measuring and improving recall in efficient language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.896786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.079998Z digest=sha256:a6c05cd88e518bcba1619fb4eb346b36ede080c77c1d05f56eb2d2276fade9cb

Observation af1324c6-0fab-4832-a706-76dc132db978 · outbound

This paper cites Fast attention requires bounded entries.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fast attention requires bounded entries

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.880702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.085462Z digest=sha256:8061bd85655a138b7e04ac0b6e148a351b658c9e6e0cc1fc248d8e76cd4d5930

Observation 8a8047c2-678c-45a3-a1ac-422c5871ea0f · outbound

This paper cites Fundamental limitations on subquadratic alternatives to transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fundamental limitations on subquadratic alternatives to transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.865671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.090223Z digest=sha256:eb0dd9ec7aab3fcf64352d7c417e33bca3271395d47ce6ed270757e10b2d7090

Observation 7945a37a-e689-418e-b8c0-acba451ad33f · outbound

This paper cites On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.850871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.095765Z digest=sha256:b8eda8e60bc52f367195ad8b28f9d97e7dbb37ed14aa59e344fd42b6378b1fd4

Observation fa65ef42-f46d-40c2-93c3-37b94a062016 · outbound

This paper cites Separations in the representational capabilities of transformers and recurrent architectures.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Separations in the representational capabilities of transformers and recurrent architectures

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.835057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.100144Z digest=sha256:a47f0374c5839bf4f3c7b2dabae183def14fc1c16d4576b15024f5fbebbc9f21

Observation 024fca3b-7758-4ed9-9ec7-c685edee9285 · outbound

This paper cites Longformer: The Long-Document Transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.104615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.104615Z digest=sha256:f2470c4d66a96b9cc84f26c2cad86ff1380751e748baa1b645d3c3982d69a21e

Observation addbc6b7-327f-4df2-a6d5-7a5383a7a3a4 · outbound

This paper cites An exploration of hierarchical attention transformers for efficient long document classification, 2022.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An exploration of hierarchical attention transformers for efficient long document classification, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.819050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.110449Z digest=sha256:eed2756c356d59e6005b4be64728348140b592c17a3489b59c4554b4a9172e11

Observation d63ca6e3-07d2-4db9-ada5-861a5bfe1370 · outbound

This paper cites Colwell, and Adrian Weller.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Colwell, and Adrian Weller

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.114975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.114975Z digest=sha256:5ca8f4b1379e00ff36fb41743ec0d695b2fdda368fda17312f850140b4e64134

Observation 208bc6d9-2b45-4ef8-adae-3785c449204f · outbound

This paper cites End-to-end object detection with transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones End-to-end object detection with transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.794430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.119593Z digest=sha256:dc481bc3ad49be78ef2a5beefc4607b039e76a7bc270c392a18faaf28d75827b

Observation c0d9153c-e51b-4936-abef-f0a471cc6c8c · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.123953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.123953Z digest=sha256:e9b0cd6b281a29a8ca01df1f451a16dd7bc8c21d96dcc9e6f231e4655df195ff

Observation 74419346-96a3-4ec1-940e-a64dd8201936 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.779890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.128702Z digest=sha256:28b2bd2a3491fa6c7391d24ddfc746245c1124d8cacd26533e2eea42916cc3e9

Observation 02e50453-d976-4a8e-a032-008066dfdcaf · outbound

This paper cites Etched is Making the Biggest Bet in AI.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Etched is Making the Biggest Bet in AI

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.765080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.133124Z digest=sha256:de4c082b5abe8e4d5c7b9d6f7687f684d55026027a3d28c28486ece6206def8c

Observation f4cb5a0f-b75d-436c-996d-6573336e49c3 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.138043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.138043Z digest=sha256:cf4c796bdd791436f2784035288f208e4730d65b39e89844fd47a483dfb2bb18

Observation a02b2ac7-3259-4954-8b3e-8da43426ddb9 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Theoretical limitations of self-attention in neural sequence models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.748865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.143413Z digest=sha256:35ac120114c1605c34f111356ec090ea9dcdbf813233d7a1d04baa597f77fe27

Observation 38ea6915-586e-4ae5-941a-ada67e86cc74 · outbound

This paper cites Hyperattention: Long-context attention in near-linear time.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hyperattention: Long-context attention in near-linear time

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.734346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.148089Z digest=sha256:3c2c361d672f5f5f61f67c0bc4a6b13351c7bde9a7840c78861590992e2abde7

Observation 08d07ec7-a333-4c40-8c90-2aa45499d470 · outbound

This paper cites Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.719060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.152253Z digest=sha256:690d42a93eb6f305e1fa04c4b04247f38f7554a21651a215435c0f35d3d0f480

Observation fe1866d1-3cbd-4e71-9b22-14369cd71c7c · outbound

This paper cites Multilayer feedforward networks are universal approximators.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Multilayer feedforward networks are universal approximators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.703811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.156366Z digest=sha256:3df897db36de7b9d97e1ef3852d3a3ab6c6356a8961e4d36098d3875911ff5c5

Observation 38e00788-d632-489c-995f-418982c0cfe1 · outbound

This paper cites On statistical rates and provably efficient criteria of latent diffusion transformers (dits).

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On statistical rates and provably efficient criteria of latent diffusion transformers (dits)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.689004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.160515Z digest=sha256:1d7185855bbe878b4e6b9c4015f87645f7863c40c5c271b95c0f010bba020ca3

Observation 7cdac334-4f01-435b-8d57-51f8908b4925 · outbound

This paper cites Kakade, and Eran Malach.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Kakade, and Eran Malach

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.165303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.165303Z digest=sha256:1bc09262ab2373adf7806a92ac472c94beb6491f5245b89ae22b41cc600fe109

Observation acebf52c-beb3-4a63-ab6c-4ac01dd9cf27 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An image is worth 16x16 words: Transformers for image recognition at scale

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.663817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.170867Z digest=sha256:6f91f977d9c0d0c141c0a5a8d1ba86cb55b2d98863adbd235a0fafe334c13d9d

Observation e9d094b8-0055-414d-9fb4-a96198d7feb0 · outbound

This paper cites RNGD preview: The world's most efficient AI chip for LLM inference.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNGD preview: The world's most efficient AI chip for LLM inference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.650041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.175070Z digest=sha256:6abbb5cd9f6dc8e62caeca3b05929da0178728fbb65652fff1aa968a642c48a2

Observation d80f281d-47c8-4a70-a344-5a40cd0df65d · outbound

This paper cites Reformer: The efficient transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reformer: The efficient transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.635829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.178891Z digest=sha256:c57048c9a69c1af22ee38e0c8df90db5080ff8cd67a2af7f037845549df97b28

Observation 278afdd6-e958-4ee1-906d-af3f472e816a · outbound

This paper cites Polysketchformer: fast transformers via sketching polynomial kernels.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Polysketchformer: fast transformers via sketching polynomial kernels

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.622136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.182787Z digest=sha256:46dae1fd9beed8d2b1d8137785d6a0b6841e87c1f2733377d2947a9bcb722119

Observation 48ed1d42-b864-410a-bf0b-22644b5e50be · outbound

This paper cites Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.608236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.186811Z digest=sha256:c33adb8c37f831a844ce7442a5d37768c4c16acb9a8c1b2e8d466cdf6cce120b

Observation ed2960ec-a1f8-4d69-88f3-f2648a08b744 · outbound

This paper cites On the expressive flexibility of self-attention matrices.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the expressive flexibility of self-attention matrices

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.593022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.190765Z digest=sha256:e6d4b833c3260b4746fb59039d8e6474bd94f0ea6cc423cbb8aaa5703c917f84

Observation c29f4383-d564-47e4-b9a4-2b207dacee37 · outbound

This paper cites Hierarchical transformers for multi-document summarization.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for multi-document summarization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.578990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.194683Z digest=sha256:5cf8b54a24c2145a45d970ac7a86036b3767ca684696f454a3c7a3d5bf63a3b1

Observation d90d4e61-9a47-4870-896a-c408cef07b8b · outbound

This paper cites The parallelism tradeoff: Limitations of log-precision transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones The parallelism tradeoff: Limitations of log-precision transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.564989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.198545Z digest=sha256:8c5b03286aac3ca79cee7be65be5e2176507d61cf743803416ca2efbd8cc2b17

Observation fc1f91c6-6d9a-44ea-9d52-57cf35de5df0 · outbound

This paper cites an unresolved cited work.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:12:48.550241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.202818Z digest=sha256:91fedf00a2d6805a2b290fa05bfbf34c17536f98c88d81286f623d88c77b0304

Observation 3bb1e02f-d1fa-4f28-9bbd-3aa12a639221 · outbound

This paper cites Language models are few-shot learners.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Language models are few-shot learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.535978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.206861Z digest=sha256:2743d91f875566bd362892bd759a96ae586ca8a97dc451365346c1d108093b67

Observation 3b258ad8-d2ad-446f-badb-12896773279a · outbound

This paper cites Hierarchical transformers for long document classification.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for long document classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.521713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.210920Z digest=sha256:e2d6a36abb77ecf5ef92956ecf072a94a39d171d7942ce86edea11ac50d8480e

Observation b661e8a2-a0a8-4843-add4-4105573376c4 · outbound

This paper cites Representational strengths and limitations of transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Representational strengths and limitations of transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.507820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.214838Z digest=sha256:996eec4c1cb92d30e81f1f070983cbfd50331e207d3d4205fe7458a112b39c8c

Observation 3ef4a2f6-b742-4a37-a5c6-ff8c5eba978d · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Transformers, parallel computation, and logarithmic depth

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.493360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.218877Z digest=sha256:077e01566dc71b0bd9af2f55f43fd07cf7ff877e08e7f0c3a1dd0ed4d3f698e0

Observation 44e1558a-3115-4c57-9557-a69b1b43a757 · outbound

This paper cites What formal languages can transformers express? a survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones What formal languages can transformers express? a survey

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.477544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.223004Z digest=sha256:18e7e0bfbab638795a85f618c8c5c786f08d28f53adb3ddcd67cbde0cb49e568

Observation c1d55e78-58a6-4b59-9ea1-a56d3a76e0b8 · outbound

This paper cites Efficient transformers: A survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient transformers: A survey

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.462466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.227406Z digest=sha256:ce5fe559f4d21d1b4cd27cc3b64912adbc7d99ae227347eaa5dfacc84075fc59

Observation 9445936f-2977-44c6-af2d-d60cdfd6cd04 · outbound

This paper cites Schmidt, and Stephan Peitz.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Schmidt, and Stephan Peitz

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.446971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.231424Z digest=sha256:e2d02f7f739cb8eb55c334715c78b4b10cfe3b963a51d6a4bc65db5cc295f36d

Observation b3536e4d-94ec-4bcd-bfec-8e67154f4f8e · outbound

This paper cites Gomez, ukasz Kaiser, and Illia Polosukhin.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Gomez, ukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.429843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.235774Z digest=sha256:8955ee5d98c4ab7ed2a7ac0daeb21408930a9d5b8a8b01bc753e25bdd3821ef2

Observation 80d2ebdf-30d9-43e7-b2e9-9935b9e57e94 · outbound

This paper cites RNN s are not transformers (yet): The key bottleneck on in-context retrieval.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNN s are not transformers (yet): The key bottleneck on in-context retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.414278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.240370Z digest=sha256:1f87a96227b286f6cdadc691fb01b9e60467b5a9845b850a7c45f57dcf332aca

Observation de93e740-306e-4720-b64c-b46d5a14bddd · outbound

This paper cites LightSeq2: Accelerated Training for Transformer-based Models on GPUs.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones LightSeq2: Accelerated Training for Transformer-based Models on GPUs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T01:12:48.303846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.245834Z digest=sha256:837e9bab7f5e5adbb8a021d86ff1b3d2652e00525c752a071654657c0a2680d8

Observation 7d215dfb-3d60-43fe-8962-399d38217616 · outbound

This paper cites Efficient streaming language models with attention sinks.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient streaming language models with attention sinks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.398365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.251148Z digest=sha256:0ba04c4a93748524c982f8b1b5ca7eb78b10b270ac2b0545afbcfa0287f91f95

Observation 12d921ac-cc04-4481-b458-297b95299f53 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.382813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.255183Z digest=sha256:7046efe9d3f4ce9ab0fead64d51fa81c8b1ad7e013c2aa805c6fdfe66b203ba2

Observation c1d80b37-089f-44c1-868f-5cd1ac530dc0 · outbound

This paper cites Reddi, and Sanjiv Kumar.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reddi, and Sanjiv Kumar

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.366538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.259338Z digest=sha256:3879ceaeba2e9fe3d3d42a387fc1773337eaa04e9495d2d237fe9459726ec973

Pith citing papers

Observation 795242bf-3d0f-49a8-96ad-9f827583f903 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Two Heads Are Better than One: Simulating Large Transformers with Small Ones

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.232597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:fd223230d943465c2b16637d74d92de47c71e567c0ce3b8325bd3a70f9c88f30