Pith. sign in

Paper Citation Record · LEDGER

Leaner Transformers: More Heads, Less Depth

As of 22 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2505.20802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20802 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:13.305354Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa05ce8e-8208-4cdb-91a9-9f0d85dfee56 · outbound

This paper cites A deep conditioning treatment of neural networks.

Leaner Transformers: More Heads, Less Depth A deep conditioning treatment of neural networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:19.425180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:09.034328Z digest=sha256:0fbad11e6bb61c3b333687ca029173fc20533b3337094d6fc25977f8d9d6e723

Observation 01750a21-676d-4050-b7cd-1a2ce70de4ec · outbound

This paper cites Xcit: Cross-covariance image transformers.

Leaner Transformers: More Heads, Less Depth Xcit: Cross-covariance image transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:19.156179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:09.162406Z digest=sha256:7920540d22a98b4ede7afdebdf89d479394fbfae6121a0edb0d58b914b89918e

Observation 622a335c-181c-42a2-be4b-66791f259c5f · outbound

This paper cites On the op- timization of deep networks: Implicit acceleration by over- parameterization.

Leaner Transformers: More Heads, Less Depth On the op- timization of deep networks: Implicit acceleration by over- parameterization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.867260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:09.316954Z digest=sha256:08a8b20658a7a38f72b2b09f7a35e81a46e757ed415d7e62a8c20b94e3f3e3a3

Observation 3d4441ca-1002-4665-9104-6606be01699b · outbound

This paper cites End-to- end object detection with transformers.

Leaner Transformers: More Heads, Less Depth End-to- end object detection with transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:09.459325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:09.459325Z digest=sha256:b51b95772842a7d26ec75f4e1351d43440efe6f353fb93d3662e93944f50a2e0

Observation 832261df-1c02-4e3d-9e0e-6dc5476c78b7 · outbound

This paper cites Rethinking attention with performers.

Leaner Transformers: More Heads, Less Depth Rethinking attention with performers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.721286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:09.561350Z digest=sha256:3e9bf11dfc6322d6e2120de4adc87e3ba66233d05a87540bf3e445ed0534032e

Observation a1419588-3169-40f8-ad3a-796f240d693b · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Leaner Transformers: More Heads, Less Depth BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:09.658901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:09.658901Z digest=sha256:4ebf45c1b127d5cdb96f88d900a74d4e8bd4e04a2c50a06025a968bfe058ddd8

Observation 980b9ea1-4735-4fe4-986c-45f30d8d05f2 · outbound

This paper cites Davit: Dual attention vision transform- ers.

Leaner Transformers: More Heads, Less Depth Davit: Dual attention vision transform- ers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.551570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:09.807743Z digest=sha256:6ced441cdd4f60bfbf7873286b6ee04b52e448887eaab8b0a6aa583a4d4ec04d

Observation 93c57150-92b7-43bf-bdaf-4e2adab4804a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Leaner Transformers: More Heads, Less Depth An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:09.946194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:09.946194Z digest=sha256:18131cca7fea63ddbe01c1cd21cf307382fa43e20d84208763fb465ae2dac360

Observation ef95e9bf-f52e-4ea4-ad23-0d8c634941a0 · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

Leaner Transformers: More Heads, Less Depth TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.115483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.115483Z digest=sha256:9524c8bd9a35462c0e4773afcf5ac70ed8a1c2d9794be4298edd7919741761f1

Observation c2f8fbce-d0ed-46bb-ae0b-cc288c2dc65a · outbound

This paper cites Drive like a human: Rethinking au- tonomous driving with large language models.

Leaner Transformers: More Heads, Less Depth Drive like a human: Rethinking au- tonomous driving with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.332722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:10.232092Z digest=sha256:2f4aa59d3a90bed6bb11637b213f89a4847a6dc9af32fd3f8da993cc19a7202c

Observation 4076683d-2bb3-46c1-844a-813c4778f542 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Leaner Transformers: More Heads, Less Depth The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.351485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.351485Z digest=sha256:5f9719c88f4842388017f2f78bd13be622c58b67bf1ab1fe21239f4ed8ae169e

Observation 94949c45-019d-4f6a-8097-fccfe740ac80 · outbound

This paper cites Cramming.

Leaner Transformers: More Heads, Less Depth Cramming

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.181276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:10.468447Z digest=sha256:43ce8f61c9147c3bdb9ea159905ada32015ca96b520091160865f6eab7b542c0

Observation 8122e36f-47df-4a41-b51d-88980b3c09d1 · outbound

This paper cites Cramming: Training a language model on a single gpu in one day.

Leaner Transformers: More Heads, Less Depth Cramming: Training a language model on a single gpu in one day

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.003646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:10.536029Z digest=sha256:11285ad7ca8d9695737921897e71b841dac84aec2385dbfb5be74a6492d3072a

Observation d44dec49-3d12-4736-8a91-3f0a003c6f62 · outbound

This paper cites Transformer in transformer.

Leaner Transformers: More Heads, Less Depth Transformer in transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.628588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.628588Z digest=sha256:7be9638d571f7a2b28bd555a415282ac33d304e5c8208a05f8c6d43174a63779

Observation a817ca96-d9b8-4b60-b04e-993a1772915f · outbound

This paper cites Neu- ral tangent kernel: Convergence and generalization in neural networks.

Leaner Transformers: More Heads, Less Depth Neu- ral tangent kernel: Convergence and generalization in neural networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.755706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:10.710054Z digest=sha256:33f62d5dd79c1418c4db6ff38facb85355a184de52ae86c3bc45f3aa063cb043

Observation d4433458-c758-451a-b672-a7cd124caa67 · outbound

This paper cites On the size of convolutional neural networks and generalization per- formance.

Leaner Transformers: More Heads, Less Depth On the size of convolutional neural networks and generalization per- formance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.540790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:10.798388Z digest=sha256:4521b4486d2b4ae455d1b8c55816d804b864b8f087a6de0457d96809d6063b72

Observation 28f6b77f-0022-4929-9d0e-ccb0d27f4a67 · outbound

This paper cites Re- former: The efficient transformer.

Leaner Transformers: More Heads, Less Depth Re- former: The efficient transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.374388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:10.880408Z digest=sha256:7a82187aa25aaf1dfe66e5600ca09ac333c16249f92b9990c8d3778615a3c4fa

Observation 8c440c53-4fd6-4377-88bc-30bded8e3cc9 · outbound

This paper cites The Depth-to-Width Interplay in Self-Attention.

Leaner Transformers: More Heads, Less Depth The Depth-to-Width Interplay in Self-Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.980819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.980819Z digest=sha256:8691cbb74fe5e21e60057a50c36d184748f8eb50d8e4687cbccdf804e53da7a8

Observation a4656912-f2d7-4c95-9724-11c96f5988fe · outbound

This paper cites Limits to depth efficiencies of self-attention.

Leaner Transformers: More Heads, Less Depth Limits to depth efficiencies of self-attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.164882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.067802Z digest=sha256:c6b4669fcd6ff9618c52059b27cd9f9653f01820f658196b938a9e94741c5022

Observation 8bcb3d35-6155-4a67-adfa-a7f17636c20f · outbound

This paper cites On Tighter Generalization Bound for Deep Neural Networks: CNNs, ResNets, and Beyond.

Leaner Transformers: More Heads, Less Depth On Tighter Generalization Bound for Deep Neural Networks: CNNs, ResNets, and Beyond

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:11.154002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:11.154002Z digest=sha256:3778bc1aef245ed001608bab75c927ab31cb3c93711c446364985d2b491b225b

Observation 424b5107-c515-41ca-ab89-a6e7104cf9e4 · outbound

This paper cites Loss land- scapes and optimization in over-parameterized non-linear systems and neural networks.

Leaner Transformers: More Heads, Less Depth Loss land- scapes and optimization in over-parameterized non-linear systems and neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.915596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.267568Z digest=sha256:6d9eb1b2878ec2718a2f1bef1e342a4166ead689f35a8d881e2b28292473929b

Observation 4b3b603c-76cd-40f3-89ff-27477b4cdb5b · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Leaner Transformers: More Heads, Less Depth Swin transformer: Hierarchical vision transformer using shifted windows

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:11.348724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:11.348724Z digest=sha256:463d434d682cdba61fe9315c01cb907bbd21fef703bb2e355ad22e849b0b34b2

Observation ed0b4e89-b93e-4d38-b190-339e79bcc449 · outbound

This paper cites The expressive power of neural networks: A view from the width.

Leaner Transformers: More Heads, Less Depth The expressive power of neural networks: A view from the width

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.716353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.438467Z digest=sha256:3e5e9221cf7b19ff5c81d3f43b491785da1101e11a672606d8749b4fb18a8ac9

Observation 721e7453-42f3-4eeb-9e84-7b01e24dea3b · outbound

This paper cites Transfusion: Multi-modal fusion network for semantic segmentation.

Leaner Transformers: More Heads, Less Depth Transfusion: Multi-modal fusion network for semantic segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.571609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.544733Z digest=sha256:6246f7c3a8a787097d5bf283fcdef3f93c303b356083c033d83bf2ec9a5ef3ef

Observation 8bf7a85e-954d-44b9-934e-6add32e88502 · outbound

This paper cites Numerical optimiza- tion.

Leaner Transformers: More Heads, Less Depth Numerical optimiza- tion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.426135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.625386Z digest=sha256:6d849acefd65f27687956c998e1060a4c1a95373e42993e0265573e6f4ae181c

Observation afafbf8c-2619-4c9a-be87-d9271319f90d · outbound

This paper cites The impact of depth and width on transformer language model generalization.

Leaner Transformers: More Heads, Less Depth The impact of depth and width on transformer language model generalization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.267841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.758211Z digest=sha256:d153f2f611ba7007ee22827193a5badd64fabf0af163ed7d94ce316c35827b76

Observation fc66014d-698a-487f-b293-a4faf56edf7e · outbound

This paper cites Exponential expressivity in deep neural networks through transient chaos.

Leaner Transformers: More Heads, Less Depth Exponential expressivity in deep neural networks through transient chaos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.130245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.859104Z digest=sha256:e737c019de08cfe4339dc897102d2424cda6b506699dd3240eaaa17ce776b733

Observation adb98fdc-43f3-4d81-bbe6-48334022aafa · outbound

This paper cites Tiny-stories-gpt.

Leaner Transformers: More Heads, Less Depth Tiny-stories-gpt

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.929889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:11.950438Z digest=sha256:e50e4f1023621d9c192768bc5792cb9c78e12e6c292493899d1403f094649b9b

Observation 4e46a6c6-03c5-4f7f-a075-8794bc32805f · outbound

This paper cites Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data.

Leaner Transformers: More Heads, Less Depth Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.040288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.040288Z digest=sha256:97a34ea6c04e113497dad1bb390538e161f221b7bad9bf252e7e1d5275b28d4c

Observation 43510b91-c664-4914-85d7-8c226607df28 · outbound

This paper cites Representational strengths and limitations of transformers.

Leaner Transformers: More Heads, Less Depth Representational strengths and limitations of transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.767405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.116907Z digest=sha256:9bfc94b280e788b39121510f0b9b2a644bf3a85d4cb3d4138f2cfcd55cb1bc17

Observation 1985930e-d536-4375-a2b6-5680575faf4e · outbound

This paper cites Real analysis: measure theory, integration, and Hilbert spaces.

Leaner Transformers: More Heads, Less Depth Real analysis: measure theory, integration, and Hilbert spaces

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.626931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.186595Z digest=sha256:a8339624f21ce5009c9f3b7d1c748ccaa116b80b00d798651813300e1d4ce9af

Observation 764911ee-d831-4bb5-84c0-49aa58cfe182 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

Leaner Transformers: More Heads, Less Depth How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.281412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.281412Z digest=sha256:a77135ea02be7dd7b78fe3f9d908f76f89d06d5d28d02ff70159ca2b9d3c8997

Observation 31d27f5e-ddd2-4935-8e67-82d97cfd151a · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

Leaner Transformers: More Heads, Less Depth Long Range Arena: A Benchmark for Efficient Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.361815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.361815Z digest=sha256:78a2ea142de91247c823e8298f5eddcb8f0d58d91bd5f23d24ff8d57906551b0

Observation 36dc04fe-54e3-4747-b9aa-194264362edc · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

Leaner Transformers: More Heads, Less Depth Training data-efficient image transformers & distillation through at- tention

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.394585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.442508Z digest=sha256:3c285e89c316737811695bcbe45afa7a50cc14433fc35d030ce1dddf2c7d43e2

Observation 5849d952-ba2b-4f5b-be49-8e315ee16638 · outbound

This paper cites Width is less important than depth in relu neural networks.

Leaner Transformers: More Heads, Less Depth Width is less important than depth in relu neural networks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.177550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.535662Z digest=sha256:202ab1aa86ad96f142618865625b513bd16a3f4d82da6f6cc10cff96d92e7e43

Observation 4244fd36-db2f-401f-9ad1-5a334b7590c3 · outbound

This paper cites Attention is all you need.

Leaner Transformers: More Heads, Less Depth Attention is all you need

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.854796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.619327Z digest=sha256:1ac88dc64e6f7c183c1c88a0ade2a937421a0e8bd48e8180547bd41dd19be586

Observation 80c06c48-fc80-4a8e-8931-33c6daf085a9 · outbound

This paper cites High-dimensional probability: An intro- duction with applications in data science.

Leaner Transformers: More Heads, Less Depth High-dimensional probability: An intro- duction with applications in data science

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.720110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.713775Z digest=sha256:ef9e7312232b0e4ccd864d267a1c77392a4415423d0bd5d5530b861f169a58ef

Observation 27206a34-a8a9-450f-aa85-cd59e28c352d · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Leaner Transformers: More Heads, Less Depth GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.781563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.781563Z digest=sha256:721ed2e8a1ddd122fabea3194b4e714fbca2c023893a6db9c5bc0ea1d21afc8c

Observation 1b162034-64cc-4bd9-bcea-25ad16772ef1 · outbound

This paper cites Linformer: Self-attention with linear complex- ity.

Leaner Transformers: More Heads, Less Depth Linformer: Self-attention with linear complex- ity

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.514716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.836097Z digest=sha256:2b0fb45d8b3c4afbdb476ebd940ddb86d303c2d655b4d727eb9b9ecb9a7c713b

Observation d50ddb4e-77f8-4be6-9e21-843ccd74a516 · outbound

This paper cites Github repository, 2021.

Leaner Transformers: More Heads, Less Depth Github repository, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.325590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.907534Z digest=sha256:759ec461b947194c50c79dca4114d4ed649ba77c1fca0fb4dc340b5fac2e586b

Observation 80d8fcc6-bb68-4c83-986a-3ff2466558af · outbound

This paper cites Nystr¨omformer: A nystr¨om-based algorithm for approximat- ing self-attention.

Leaner Transformers: More Heads, Less Depth Nystr¨omformer: A nystr¨om-based algorithm for approximat- ing self-attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.156866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:12.982547Z digest=sha256:5f3fd63696597f8250b70bb976f1e8ffdf4af894a875a1db062a25a85e317849

Observation 622798dd-8cab-497c-95d4-262d39936b60 · outbound

This paper cites V olo: Vision outlooker for visual recog- nition.

Leaner Transformers: More Heads, Less Depth V olo: Vision outlooker for visual recog- nition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.014039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:13.044588Z digest=sha256:d52eb4dd3cb6237b90ae5fa824d4037c8550a52beafe46e8e7fd76756663ccd7

Observation 73f1affc-3c99-4491-a8f3-2d1395b2a604 · outbound

This paper cites cosformer: rethinking softmax in attention.

Leaner Transformers: More Heads, Less Depth cosformer: rethinking softmax in attention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:13.848094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:13.118002Z digest=sha256:2be93ddf40eeaf289f088371d012c73da451dd2bb4830bc15ba37b7e8b36e619

Observation 2d421d34-2722-4ebe-b59a-3ee49888125a · outbound

This paper cites Understanding generalization and optimization performance of deep cnns.

Leaner Transformers: More Heads, Less Depth Understanding generalization and optimization performance of deep cnns

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:13.683260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:13.197390Z digest=sha256:e9e6236b7f436c78538d8edc1c58ffba1146839d8839549ab5c5253087f7e5f4

Observation edd4c23a-d92b-4435-82b4-51a4a181152e · outbound

This paper cites A robustly optimized bert pre-training approach with post-training.

Leaner Transformers: More Heads, Less Depth A robustly optimized bert pre-training approach with post-training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:13.523878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:53:13.305354Z digest=sha256:40f1f61b2bc29c42ad5c1a2b2597eb3bf992d9a9e0d2aff804b7d6b71676cc50

Pith citing papers

No inbound Pith citation observations are available.