Pith. sign in

Paper Citation Record · LEDGER

Two Heads Are Better than One: Simulating Large Transformers with Small Ones

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12220 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:12:48.259338Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:05:58.230449Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b51397b-9819-4a56-bba0-0277ff4eb8cd · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Zoology: Measuring and improving recall in efficient language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.896786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.079998Z digest=sha256:a0457eeae809dd832d10b493ecd58f2f61d11f061879379616ffd6a13dd35fad

Observation af1324c6-0fab-4832-a706-76dc132db978 · outbound

This paper cites Fast attention requires bounded entries.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fast attention requires bounded entries

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.880702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.085462Z digest=sha256:7f3bfba6996a4855af79eed27e6f4955e69a795020095d8f6cce8bff1a0e13e6

Observation 8a8047c2-678c-45a3-a1ac-422c5871ea0f · outbound

This paper cites Fundamental limitations on subquadratic alternatives to transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fundamental limitations on subquadratic alternatives to transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.865671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.090223Z digest=sha256:3db60ba622bdf7175cf9a3ba083ad8bb45c5f62813a19a4e46a30eaab1913257

Observation 7945a37a-e689-418e-b8c0-acba451ad33f · outbound

This paper cites On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.850871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.095765Z digest=sha256:ca9f63553461e789a561e544fe41b854cfed899ce06a32a1888f633246a32a9c

Observation fa65ef42-f46d-40c2-93c3-37b94a062016 · outbound

This paper cites Separations in the representational capabilities of transformers and recurrent architectures.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Separations in the representational capabilities of transformers and recurrent architectures

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.835057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.100144Z digest=sha256:d468ebbe064c70f185edb885ea3457624a494896197f7576a1f1f4a20b99746a

Observation 024fca3b-7758-4ed9-9ec7-c685edee9285 · outbound

This paper cites Longformer: The Long-Document Transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.104615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.104615Z digest=sha256:7dfcc7f9d4731d4914ff6746ded3025af8e939f343fce2afa40d6dc8c051122d

Observation addbc6b7-327f-4df2-a6d5-7a5383a7a3a4 · outbound

This paper cites An exploration of hierarchical attention transformers for efficient long document classification, 2022.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An exploration of hierarchical attention transformers for efficient long document classification, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.819050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.110449Z digest=sha256:cce165fdb7ce110311c025bfd7c3664956faa45be7cca164092768209c7c668f

Observation d63ca6e3-07d2-4db9-ada5-861a5bfe1370 · outbound

This paper cites Colwell, and Adrian Weller.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Colwell, and Adrian Weller

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.114975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.114975Z digest=sha256:61ba559fb1b7cfc1f08cef71a99859b41a820d0eed32d5f8932bf451ccff837e

Observation 208bc6d9-2b45-4ef8-adae-3785c449204f · outbound

This paper cites End-to-end object detection with transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones End-to-end object detection with transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.794430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.119593Z digest=sha256:ec8d071efe94e4bb8c5f573ce8b6daf45a5bdd020552222b718d001712f9aa4e

Observation c0d9153c-e51b-4936-abef-f0a471cc6c8c · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.123953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.123953Z digest=sha256:9d5b2f5dc2ba26047eff0302723666eb273cc9e8361dfa48dae32ff0b52ef4b2

Observation 74419346-96a3-4ec1-940e-a64dd8201936 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.779890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.128702Z digest=sha256:f8a37c259a14844cd8a4c0e333194ce56ec0eeafcc85169c0d5748c71c59077e

Observation 02e50453-d976-4a8e-a032-008066dfdcaf · outbound

This paper cites Etched is Making the Biggest Bet in AI.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Etched is Making the Biggest Bet in AI

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.765080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.133124Z digest=sha256:923b712e0cd514c4dfafbca4015de46a09443e08041ffea13e7213d0a065a54c

Observation f4cb5a0f-b75d-436c-996d-6573336e49c3 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.138043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.138043Z digest=sha256:d015067561438fa013d06cd49b54b82d669901abee803d715ba0fbd2063a33d1

Observation a02b2ac7-3259-4954-8b3e-8da43426ddb9 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Theoretical limitations of self-attention in neural sequence models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.748865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.143413Z digest=sha256:9f8f4d5332a07ff7f2a736ed57baf69d90501e4dcdaa55e9a7a3a5d6a982f9f7

Observation 38ea6915-586e-4ae5-941a-ada67e86cc74 · outbound

This paper cites Hyperattention: Long-context attention in near-linear time.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hyperattention: Long-context attention in near-linear time

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.734346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.148089Z digest=sha256:0c8ea19be0d9a9c31223a9cfe5bcde622ae52156aca360a1932f67051d63feb9

Observation 08d07ec7-a333-4c40-8c90-2aa45499d470 · outbound

This paper cites Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.719060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.152253Z digest=sha256:bcf7960c8b4d65d6c1343bf19573fc823e790660c133dcae1de48c8969639f0a

Observation fe1866d1-3cbd-4e71-9b22-14369cd71c7c · outbound

This paper cites Multilayer feedforward networks are universal approximators.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Multilayer feedforward networks are universal approximators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.703811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.156366Z digest=sha256:9253629382a014a0d9365d3fee4f904c25bd4272e9afdf4cfe8c426815877a13

Observation 38e00788-d632-489c-995f-418982c0cfe1 · outbound

This paper cites On statistical rates and provably efficient criteria of latent diffusion transformers (dits).

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On statistical rates and provably efficient criteria of latent diffusion transformers (dits)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.689004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.160515Z digest=sha256:612dc75bac0339306475722797675cf8928e35b1ac97ccac13fc6507918a2cb5

Observation 7cdac334-4f01-435b-8d57-51f8908b4925 · outbound

This paper cites Kakade, and Eran Malach.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Kakade, and Eran Malach

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.165303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.165303Z digest=sha256:61166a0412f835ecd1f5b540bb3cfe51ba6931b910e338d31a8df7ca2cad945b

Observation acebf52c-beb3-4a63-ab6c-4ac01dd9cf27 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An image is worth 16x16 words: Transformers for image recognition at scale

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.663817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.170867Z digest=sha256:aea20f7f4280e3eb18d4050acef08da5d1b16cd64fad121f9c1ca8f82c904fd0

Observation e9d094b8-0055-414d-9fb4-a96198d7feb0 · outbound

This paper cites RNGD preview: The world's most efficient AI chip for LLM inference.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNGD preview: The world's most efficient AI chip for LLM inference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.650041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.175070Z digest=sha256:76dd2fc53845dfe2c9aaece7832c15a70d2db0b2a1f43fa025526b636c368c6c

Observation d80f281d-47c8-4a70-a344-5a40cd0df65d · outbound

This paper cites Reformer: The efficient transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reformer: The efficient transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.635829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.178891Z digest=sha256:f4b30333c396ce01c2ebc540f2679df4313ff3b0f464a98c141621a79247a5aa

Observation 278afdd6-e958-4ee1-906d-af3f472e816a · outbound

This paper cites Polysketchformer: fast transformers via sketching polynomial kernels.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Polysketchformer: fast transformers via sketching polynomial kernels

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.622136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.182787Z digest=sha256:ae34e3bd3d0953b95d4a107a0487a3f61fc2c7adeaacf2554f086a7e17b70fe5

Observation 48ed1d42-b864-410a-bf0b-22644b5e50be · outbound

This paper cites Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.608236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.186811Z digest=sha256:b7fc7b82ec79b3deb677c902403687c8779f332e643edd91ed327355a9a5dfd6

Observation ed2960ec-a1f8-4d69-88f3-f2648a08b744 · outbound

This paper cites On the expressive flexibility of self-attention matrices.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the expressive flexibility of self-attention matrices

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.593022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.190765Z digest=sha256:565d11adff1fa95275fd1a4101e64ea18ed09320a9157e59e836deef253ae288

Observation c29f4383-d564-47e4-b9a4-2b207dacee37 · outbound

This paper cites Hierarchical transformers for multi-document summarization.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for multi-document summarization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.578990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.194683Z digest=sha256:58a058a8c6d02a8fd91f92a9a284e8b3e3c254cacb1fb9627a408683e2897ef5

Observation d90d4e61-9a47-4870-896a-c408cef07b8b · outbound

This paper cites The parallelism tradeoff: Limitations of log-precision transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones The parallelism tradeoff: Limitations of log-precision transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.564989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.198545Z digest=sha256:b93546b7ed946426a0712c4c10a1c303dcb9722d84c431cef47176250891ae80

Observation fc1f91c6-6d9a-44ea-9d52-57cf35de5df0 · outbound

This paper cites an unresolved cited work.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:12:48.550241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.202818Z digest=sha256:cb09179b0b95abfe0edc3eb445215772b002797e232640d5b899acf75b4ad3ac

Observation 3bb1e02f-d1fa-4f28-9bbd-3aa12a639221 · outbound

This paper cites Language models are few-shot learners.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Language models are few-shot learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.535978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.206861Z digest=sha256:3e5de6e21177784a7f80ae149b70305c6425606628007761ccc7e257b44fa5f6

Observation 3b258ad8-d2ad-446f-badb-12896773279a · outbound

This paper cites Hierarchical transformers for long document classification.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for long document classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.521713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.210920Z digest=sha256:db0fbd2494a6f1742a1a1686b5dfe4a5514a2d2cc8beff660475070d1c37180d

Observation b661e8a2-a0a8-4843-add4-4105573376c4 · outbound

This paper cites Representational strengths and limitations of transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Representational strengths and limitations of transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.507820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.214838Z digest=sha256:1416b183c5b940d3fea0f78356a00f97f24de090021a45c8a266be551a844dcf

Observation 3ef4a2f6-b742-4a37-a5c6-ff8c5eba978d · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Transformers, parallel computation, and logarithmic depth

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.493360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.218877Z digest=sha256:9df4b372083c19b4b255408e5fb21d080fcd75cd2efd5d88ab121b800edb4864

Observation 44e1558a-3115-4c57-9557-a69b1b43a757 · outbound

This paper cites What formal languages can transformers express? a survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones What formal languages can transformers express? a survey

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.477544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.223004Z digest=sha256:5b646b1b0261700b276b53a4fb5931b38bce6c3b7e2acc9fcd7c0e679550831a

Observation c1d55e78-58a6-4b59-9ea1-a56d3a76e0b8 · outbound

This paper cites Efficient transformers: A survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient transformers: A survey

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.462466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.227406Z digest=sha256:41b73b1b379f160636069898e32c3c95edb5c94a8e4ccb1094bbb83caee9a811

Observation 9445936f-2977-44c6-af2d-d60cdfd6cd04 · outbound

This paper cites Schmidt, and Stephan Peitz.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Schmidt, and Stephan Peitz

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.446971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.231424Z digest=sha256:095ea2d6df5079a8b93403cdc9910b0576b7529f5043a3d6b7515cd3cd292161

Observation b3536e4d-94ec-4bcd-bfec-8e67154f4f8e · outbound

This paper cites Gomez, ukasz Kaiser, and Illia Polosukhin.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Gomez, ukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.429843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.235774Z digest=sha256:2a4629f93ec2d6fabaeea2c85dc5666593e152397034406c2416ec63a73e944f

Observation 80d2ebdf-30d9-43e7-b2e9-9935b9e57e94 · outbound

This paper cites RNN s are not transformers (yet): The key bottleneck on in-context retrieval.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNN s are not transformers (yet): The key bottleneck on in-context retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.414278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.240370Z digest=sha256:da6cd803cae56926a515ee6994112eae75b3daf9a50c36e3f798bbcd5d0d4735

Observation de93e740-306e-4720-b64c-b46d5a14bddd · outbound

This paper cites LightSeq2: Accelerated Training for Transformer-based Models on GPUs.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones LightSeq2: Accelerated Training for Transformer-based Models on GPUs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T01:12:48.303846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.245834Z digest=sha256:33738a2dd5d0ff410255fe20201b635158e5538640c1c20667507d5fd3415194

Observation 7d215dfb-3d60-43fe-8962-399d38217616 · outbound

This paper cites Efficient streaming language models with attention sinks.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient streaming language models with attention sinks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.398365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.251148Z digest=sha256:3fb0bf49d9d72b61225db9fc936537a3bb9b3fe4e231db6ed1e148ea5c1a47d7

Observation 12d921ac-cc04-4481-b458-297b95299f53 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.382813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.255183Z digest=sha256:c66aafccf660d54963d2c4b01e051374530dc9eb76dacd8d640f2e29d896f7b2

Observation c1d80b37-089f-44c1-868f-5cd1ac530dc0 · outbound

This paper cites Reddi, and Sanjiv Kumar.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reddi, and Sanjiv Kumar

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.366538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.259338Z digest=sha256:fab5175a3cb0b2f43c2e9a408d65ed6cbfc9ef43d0de3571575fb454e537eb39

Pith citing papers

Observation 795242bf-3d0f-49a8-96ad-9f827583f903 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Two Heads Are Better than One: Simulating Large Transformers with Small Ones

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.232597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:fc6450b2c13e23b0c7bb27b8c5b42de6599982afdcd31c62cab44b74fe351bd4