Pith. sign in

Paper Citation Record · LEDGER

Primer: Searching for Efficient Transformers for Language Modeling

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2109.08668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.08668 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:04:38.115083Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:14:01.480491Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1e1eadf9-6752-4c38-8a22-e808dfd44f84 · inbound

ST-MoE: Designing Stable and Transferable Sparse Expert Models cites this paper.

ST-MoE: Designing Stable and Transferable Sparse Expert Models Primer: Searching for Efficient Transformers for Language Modeling

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:14:25.861443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T23:14:25.431471Z digest=sha256:8d9c076d4533f299692656cddf817aff01e81b90ee2947fb4e11e2adc0f69364

Observation 152da4b4-8478-4165-8a0c-153ba16efdd4 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Primer: Searching for Efficient Transformers for Language Modeling

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.140017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:e021989c3ab233355a83aa2d76139ed6a7960586c88aac110749766f9cd526aa

Observation 7e631a03-93fe-489a-ae0b-77f276dc3fe6 · inbound

Fast Inference from Transformers via Speculative Decoding cites this paper.

Fast Inference from Transformers via Speculative Decoding Primer: Searching for Efficient Transformers for Language Modeling

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:52:00.193939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T22:52:00.101612Z digest=sha256:ce9b8ef7bbc183777f881928acebef0f7afd44dde27917dfa28551e62a5a4aca

Observation ba483445-22e1-47fa-bc2e-426623c4cc27 · inbound

TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations cites this paper.

TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations Primer: Searching for Efficient Transformers for Language Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:38.115083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:38.115083Z digest=sha256:48dfb154f17d63ac6af6d2b33503d1f5f4121a1b81383ff5a4559bbe2b9399f5

Observation 9b6b6d0c-110e-4537-ad71-4e3d70ec4747 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Primer: Searching for Efficient Transformers for Language Modeling

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:19.008474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:19:25.483640Z digest=sha256:50ed517882547c9f9e192e6224f80d13195d4858aacff807aba536bc84803995

Observation a97283ea-c4a4-401d-a35e-43b2ee9dd29d · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Primer: Searching for Efficient Transformers for Language Modeling

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:54:51.152308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T11:54:29.436149Z digest=sha256:aba5861667f2fbac0ffe7e81b802aa8495835dc7913fd073b4a6d9a0d20d81ac

Observation 2dc3b978-431c-4f0a-9962-9642290226d5 · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Primer: Searching for Efficient Transformers for Language Modeling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:52.961943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:52.961943Z digest=sha256:1b218689d50c7fc4d141500f1b8f451eb0caf89aa4780fb124cf12858c2e1838

Observation d9ce9369-a7ab-4048-ab75-8ec05532f67d · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence Primer: Searching for Efficient Transformers for Language Modeling

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:40:42.503282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:4e42e797da21abeeb444dac7b653ecb2e4f745fe060c42f148d0e95d3d8b049e

Observation 31333250-5087-409d-9d92-d90c28a5c707 · inbound

From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums cites this paper.

From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums Primer: Searching for Efficient Transformers for Language Modeling

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:47:33.252670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:42:48.141279Z digest=sha256:b38d9a7b38d54a12bacc7804640163e11e1213a6680726f6d29cb3788d8addb6

Observation a92a4ad1-0ee7-4fac-9646-899228e1fa5e · inbound

Three-Phase Transformer cites this paper.

Three-Phase Transformer Primer: Searching for Efficient Transformers for Language Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:05:24.629483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:04:23.797546Z digest=sha256:87461a6f1e7f40859bbdf095c3ad384a37934fe6167094250ee057801989df74

Observation 81a24bda-88dc-4f43-8406-f21e617dc975 · inbound

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity cites this paper.

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity Primer: Searching for Efficient Transformers for Language Modeling

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:26:12.854349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T17:07:18.278784Z digest=sha256:9510dbc463549dcd3a961fe15d1c24083b150df9bde8701bd3da93c38b09aa68

Observation 9af5fca8-be7c-47d2-ab7e-c05513672db6 · inbound

On the global convergence of gradient descent for wide shallow models with bounded nonlinearities cites this paper.

On the global convergence of gradient descent for wide shallow models with bounded nonlinearities Primer: Searching for Efficient Transformers for Language Modeling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:31.411140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:51:08.871267Z digest=sha256:3a4b2c2060e2cd3f220b9751459fed46953436c441f91a0f505a46295f6c7d63

Observation 9e074bf5-44df-4e0c-8516-be3272f0c9cb · inbound

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes cites this paper.

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes Primer: Searching for Efficient Transformers for Language Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:48:19.650855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T13:46:32.405079Z digest=sha256:994e5f4f5be1f3264314e204e261526e28b2c7bed6b16417b6c93cbe2e6040b8

Observation 41f62084-2ea2-415e-bf29-d466206687d0 · inbound

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes cites this paper.

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes Primer: Searching for Efficient Transformers for Language Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:16:19.681494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T09:15:32.395442Z digest=sha256:05b811e0538eaff96a4079f8e242d5f3f606aadba4122c12ecb0672f1eecf2e0

Observation 03ec595f-8600-461a-b6b7-3a795559c30a · inbound

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training cites this paper.

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training Primer: Searching for Efficient Transformers for Language Modeling

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:14:01.482003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:10:47.199537Z digest=sha256:a3c36ed5ea0f4ae50d4a680a356bb30bad12e6b07a65b2cce79127b1e5ffbc5f

Observation 430ad754-b9ad-4167-82e4-4d64d919e1c6 · inbound

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks cites this paper.

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks Primer: Searching for Efficient Transformers for Language Modeling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T17:31:49.346142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:31:49.346142Z digest=sha256:bc9086da39234364cf87939543e117f879133e2ced8c1a7c595dcc439289f94b

Observation 9f1d067a-18da-46a7-b512-07118b323b7f · inbound

Domyn-Small: A European 10B Reasoning Language Model cites this paper.

Domyn-Small: A European 10B Reasoning Language Model Primer: Searching for Efficient Transformers for Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T14:07:52.153859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:07:52.153859Z digest=sha256:35a8cdaa27e45cb4b0b301027186bea8121a513db5af8c1d7671fd2b3bc95129