Pith. sign in

Paper Citation Record · LEDGER

TopK Language Models

As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21468 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:31:39.220992Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 789d64b4-69c1-41fb-80dd-51772945f630 · outbound

This paper cites How can we be so dense? the benefits of using highly sparse representations, 2019.

TopK Language Models How can we be so dense? the benefits of using highly sparse representations, 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.922624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:19.862785Z digest=sha256:536742de62c9544eb180796f4f214c9ce869e49cd39f58c6b811833e841b2fa8

Observation 252a90f8-e166-4ccf-8aa8-0427bd7175dd · outbound

This paper cites Generating long sequences with sparse transformers, 2019.

TopK Language Models Generating long sequences with sparse transformers, 2019

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.787577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.787577Z digest=sha256:ce985aaa126c485da5f5fdbed3b313072831c668715fc2fd5625626078f64a02

Observation 674b7941-d134-4ec2-b182-617c4b245cb5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

TopK Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.894657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.894657Z digest=sha256:e96b5a325972451dc68b745c2970f912ee4bcd15dcbdf9fe613d2baf82557f8c

Observation 7a6bcd59-399b-4e81-a450-405ea058ef91 · outbound

This paper cites [Full Post] Progress Update #1 from the GDM Mech Interp Team.

TopK Language Models [Full Post] Progress Update #1 from the GDM Mech Interp Team

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.790819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:35.076386Z digest=sha256:39db61750a66e1f041218f71f4cbb2201f1f7634d125cc1997192e4312c58ac1

Observation ba9b874e-ce44-4514-9642-ba67d7764f98 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models, 2023.

TopK Language Models Sparse autoencoders find highly interpretable features in language models, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.158355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.158355Z digest=sha256:3ebabf63b7e50cc6033d8fdd30a28fa6f088cac5d05e0d7064ee8bb5d9102016

Observation f0204e56-9a10-4e7c-947b-1d805f936fb0 · outbound

This paper cites Rigging the lottery: Making all tickets winners.

TopK Language Models Rigging the lottery: Making all tickets winners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.685373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:35.321207Z digest=sha256:0608548bd2c57dcb476a3fec0665c20fbf28e31f23a5e628f71d6ce63cc3cc4c

Observation 30ea9f6d-00f7-4f49-964f-0f3dc6a16879 · outbound

This paper cites Sparsity in transformers: A systematic literature review.

TopK Language Models Sparsity in transformers: A systematic literature review

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.560982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:35.434488Z digest=sha256:7a813e8bfcbf8b191e255927c3c8efb3de66388bb7f3421e92d328058fa1f7ea

Observation 72d4a652-94cd-4b34-a2aa-d89f288ff9b0 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

TopK Language Models Detecting hallucinations in large language models using semantic entropy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.530869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.530869Z digest=sha256:5c50e1212b0fcbb75dbae700a7a4c8b56c3cd70898d2ecf50acce78105b36af0

Observation 6f137638-06ec-406b-9ce0-061a3bf57b53 · outbound

This paper cites Scaling and evaluating sparse autoencoders, 2024.

TopK Language Models Scaling and evaluating sparse autoencoders, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.452327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:35.680217Z digest=sha256:b08b5a05e296e9cba5e6b41c58cfa0cdce447d03c353611f059f034b1240c431

Observation f4e0c28a-7209-4240-955b-333c6e0ea7da · outbound

This paper cites The language model evaluation harness, 2024.

TopK Language Models The language model evaluation harness, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.370392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:35.732571Z digest=sha256:f7d0f7fcc6e6d0b5f28434db22e7c6ff1716356054f3c765f84ddf31bef8615e

Observation ee4dd635-db57-410b-9324-1787e68f3a5c · outbound

This paper cites Causal Abstractions of Neural Networks.

TopK Language Models Causal Abstractions of Neural Networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.255996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:35.794053Z digest=sha256:48211c234a0e151d643ca84c8cc3215b486ca089ab8631983b3f80324e04e217

Observation 04f5fb92-8a51-49c8-b8d4-d99b82567568 · outbound

This paper cites The State of Sparse Training in Deep Reinforcement Learning.

TopK Language Models The State of Sparse Training in Deep Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.158063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:35.872432Z digest=sha256:fe80ad308d649176b74a81efd9a2904721126a4181ee13abcd2f880bda0adde1

Observation aa35ef23-a1e2-4365-8576-b6daf9a5c923 · outbound

This paper cites Memory-efficient transformers via top-k attention, 2021.

TopK Language Models Memory-efficient transformers via top-k attention, 2021

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.047952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.032941Z digest=sha256:054235e792b55d737c6d06b71bb90a5c45a354781eb91773e58be5120c9b2c30

Observation 3f7984cd-049b-4bda-86f3-9384fd31b028 · outbound

This paper cites Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024.

TopK Language Models Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.933604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.096247Z digest=sha256:c7c5a5e1e0d2b77dabcd590b98b82e42681e9665db2c2e766fd6527d2e28f4d3

Observation cdb4a400-794b-422a-b9db-23e9e5554868 · outbound

This paper cites Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba.

TopK Language Models Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.826679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.144498Z digest=sha256:ed4bc994d797e590a3f512da5756d705a49142fc8f33fc4f434d0aaa12725d01

Observation 93a23817-5ae8-4063-9a61-cc7a57c7407d · outbound

This paper cites Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks.

TopK Language Models Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:39.565510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.250242Z digest=sha256:fcabcda4f425441c831678da823f85d49c2e7503011f2a2f8e7af75b20761bee

Observation 431d5c70-f9cd-434d-b265-d0c8d173ceb6 · outbound

This paper cites How llms learn: Tracing internal representations with sparse autoencoders, 2025.

TopK Language Models How llms learn: Tracing internal representations with sparse autoencoders, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.727068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.341495Z digest=sha256:33d5211767274da6b13e0805f73916a27526df5bcde4db1dfa4126bafd9afefd

Observation d3c0bed0-dcef-4bad-a2a3-7229280fc0ad · outbound

This paper cites Sparse is enough in scaling transformers, 2021.

TopK Language Models Sparse is enough in scaling transformers, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.615792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.403525Z digest=sha256:0448d4e221d5c6e0cfe29ce91f708d223cb7151bf5607b07e224884247eacc7d

Observation d7578cb5-abb8-45ff-b7d9-85726b9abc70 · outbound

This paper cites Top-KAST: Top-K Always Sparse Training.

TopK Language Models Top-KAST: Top-K Always Sparse Training

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:31:39.411734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.509531Z digest=sha256:9808f746121bc834dcfcef1d8403b1064225c901c726d93688b99fa0d206837d

Observation 4f91072b-f53e-4f6f-b9bf-d3d8d7b64d13 · outbound

This paper cites Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025.

TopK Language Models Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.505754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.629380Z digest=sha256:6d2555374b5a8cb3a32beca837f77eb9781aa0c08cc53030c7f92a91016a7341

Observation 1d8d7c24-9783-465d-9c80-be28083f430a · outbound

This paper cites Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025.

TopK Language Models Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.327447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.737124Z digest=sha256:d77529b90f97191242b4cc29d515ba7a175ce95491a689248f9f30fa9ab787b3

Observation 88cbb3b3-8ac3-4ba0-8158-3b1251dba858 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

TopK Language Models Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.842439Z digest=sha256:4a5d7a15ce25200bc18461831cfebcb8bd3a6afbaa6dbe21634514397e09f4dc

Observation 3edfc9c7-94f6-4281-a590-880224887ad8 · outbound

This paper cites Soft Threshold Weight Reparameterization for Learnable Sparsity.

TopK Language Models Soft Threshold Weight Reparameterization for Learnable Sparsity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.002177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.925606Z digest=sha256:bab9e7f665259562cd3e64829076f42d6a06408d958091bf2d9a2ede115f95a1

Observation a05d54fd-ae7b-4413-b870-4cecc3413d54 · outbound

This paper cites Sparse autoencoders do not find canonical units of analysis, 2025.

TopK Language Models Sparse autoencoders do not find canonical units of analysis, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.807089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:36.983433Z digest=sha256:c0109181746d8caee1d58099c62f307f28a722aab144271f12fb9b147d155db1

Observation 5f4fba51-bb4a-4ebd-9a82-f17b303fec40 · outbound

This paper cites Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024.

TopK Language Models Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.056013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.056013Z digest=sha256:1b7c9bf3129e9e6b9d675ade09f794d89aab1d0d397b60be4e3990383c15fcaa

Observation 0e04cee2-46f0-46d1-ade5-282892edb508 · outbound

This paper cites Decoupled Weight Decay Regularization.

TopK Language Models Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.178870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.178870Z digest=sha256:4db6b380a820a343b856ec29cf8a5a22ec46a3aaa6aa1e2c1aba98ac522f2b43

Observation dfa20b76-4411-44f4-807f-6be32170786e · outbound

This paper cites Learning sparse neural networks through l_0 regularization.

TopK Language Models Learning sparse neural networks through l_0 regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.623753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.243550Z digest=sha256:d068d11fe1b53dd6fab7a3e16b4b136d715be56e272758db4d81babae7f4ceb7

Observation 8a792f02-b2c6-45e8-b471-bb1f4f042e41 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

TopK Language Models Fineweb-edu: the finest collection of educational content, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.336793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.336793Z digest=sha256:95ab06bf4ec18e03dbb65a0fb94bce7b2df48b9115be44785889b13482bceddc

Observation c81c18f1-a5e1-419b-878a-9a4b2a29115b · outbound

This paper cites Winner-take-all autoencoders.

TopK Language Models Winner-take-all autoencoders

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.399846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.409944Z digest=sha256:12f1880b679465ba9caafeaf63900e59cc665481a9a4067eef3a9a9bb14e55d9

Observation cf1a13db-df4e-4f85-8d87-3a1f51a45d56 · outbound

This paper cites Pointer Sentinel Mixture Models.

TopK Language Models Pointer Sentinel Mixture Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.252639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.516145Z digest=sha256:063c2c4cfb7c4d6e7e23dc3ffc5aaa989e5c13c9d6e1d3933bc2a2e96b424076

Observation 6cc3d4c0-c29f-45dc-b26e-406dba71cc93 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

TopK Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.067358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.572894Z digest=sha256:829fe7c6993ae4409595339105754f1ff76024ae4ad8c256b7681f5d5eb8591c

Observation 69e9c625-3b6b-45ee-883a-cc68e008da50 · outbound

This paper cites Nguyen, Madeleine Gibescu, Antonio Liotta, and et al.

TopK Language Models Nguyen, Madeleine Gibescu, Antonio Liotta, and et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.886349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.620117Z digest=sha256:2aa1715253c150665ab9b04d20d7c33d05122854c0932b3539fb4f983284776d

Observation 2a719d09-833e-4914-8ee3-0b74cd9203b5 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

TopK Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.688817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.666381Z digest=sha256:918a14b32dfbb6b4e0e87e933cc80a83583d5583f6ecf809b6fbf0fa6b1be24c

Observation 6dd0f983-9ab7-4e9c-8faf-5f2e127b6e35 · outbound

This paper cites Sparse autoencoders trained on the same data learn different features, 2025.

TopK Language Models Sparse autoencoders trained on the same data learn different features, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.469546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.703915Z digest=sha256:8b05a77b0e022660331b9d83153b3cf88d32747092dffccc12f6708c2716cf98

Observation 0754bfa7-a587-49c0-ab97-1f49c4ce81b1 · outbound

This paper cites Automatically interpreting millions of features in large language models, 2024.

TopK Language Models Automatically interpreting millions of features in large language models, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.268686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.798385Z digest=sha256:a4c638f12a8f8e836cf520716a6f4c62f754f0fa46fc2d5dc3a0f2dabbeafa83

Observation a1125760-f628-4ca2-9ebb-4d6f608633bf · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

TopK Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.035447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:37.854065Z digest=sha256:5f4fd3c65b329ed1a43761b8a932f674bce4fe2960e3bd11aa6da8f51de3efaf

Observation d02c0341-8289-4f6a-9e93-7c948279a759 · outbound

This paper cites Improving dictionary learning with gated sparse autoencoders, 2024.

TopK Language Models Improving dictionary learning with gated sparse autoencoders, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.972247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.972247Z digest=sha256:0be48c059cc459bcaf34aec05b681740cffcd0036c06d28067490956f04f28a7

Observation acedd645-2f42-434a-8ef1-aebff0af90bb · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale.

TopK Language Models Winogrande: an adversarial winograd schema challenge at scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.142907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.142907Z digest=sha256:d1a5e31379fd52d1650f9b058b83bb8357290609389e9cb620999e03fb7fded8

Observation 8f6dab73-abb4-4500-b8d5-f326759064a8 · outbound

This paper cites Taking features out of superposition with sparse autoencoders.

TopK Language Models Taking features out of superposition with sparse autoencoders

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.851527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:38.258973Z digest=sha256:d761bfb5db5669e0809e3d7dc9f0464d511249396712015681ae7d0dc486cb9b

Observation ed8b5f35-897e-4c8d-9f38-e00bfacbbac6 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017.

TopK Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.370614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.370614Z digest=sha256:f3c2f4e63e0a907a82d2591574edf7571cb9b8824d69dda4bc42577619a4de98

Observation 1ddad81e-ea7a-41e9-b420-c9e5d64197bd · outbound

This paper cites A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025.

TopK Language Models A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.670077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:38.509917Z digest=sha256:5d09bfd6410d3876f5d0697042b6828a3984d6d5026796262f611a1b42b7f811

Observation 9c0496bf-71fc-42ea-a76c-b215fc6f7781 · outbound

This paper cites Codebook Features: Sparse and Discrete Interpretability for Neural Networks.

TopK Language Models Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.680344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.680344Z digest=sha256:f9d220a6f8e89c3697200c6a278fe87947cc5f6ad47448f08074f46d8d1e121b

Observation e1b8015f-8dca-48ee-a6d5-94e0197c7317 · outbound

This paper cites Daniel Freeman, Theodore R.

TopK Language Models Daniel Freeman, Theodore R

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.462326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:38.794923Z digest=sha256:8eb29f0deb90a633f5877e95d3ec5dbc5ea9607fa299976956fadc4dfd7bbc34

Observation 4e902680-9768-47c0-aae7-a1a9b5f98c7c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TopK Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.890079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.890079Z digest=sha256:bc6edfade7fbdc168784747b81b7926f23987cd1b4675e804365de09a6c5520f

Observation 3ea79e63-1c97-4e4a-ad6c-80dc996c16ab · outbound

This paper cites Meta Lingua: A minimal PyTorch LLM training library, 2024.

TopK Language Models Meta Lingua: A minimal PyTorch LLM training library, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.282755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:38.998159Z digest=sha256:de29673f6c17b15778c46a96831b7b5b4087c71553999fbfd8361681adca7b20

Observation 373449a0-f1a3-45e0-b63a-24c6155ab499 · outbound

This paper cites Tracking the feature dynamics in llm training: A mechanistic study, 2025.

TopK Language Models Tracking the feature dynamics in llm training: A mechanistic study, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.093824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:39.060459Z digest=sha256:e1a714c0cedaac9ef127b18f42da25da7a442e752cd37b43fab2d54030f56daa

Observation b26d6674-e1c9-49fd-a3d2-63892221c21a · outbound

This paper cites STEP: Staged parameter-efficient pre-training for large language models.

TopK Language Models STEP: Staged parameter-efficient pre-training for large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.921354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:39.121952Z digest=sha256:9eb7681f3a70430aca2bd53a87121a3b3c6a24f428b8d07a921f628b9c2951f0

Observation 4a13895a-5f6b-42e3-a2c2-6056c77dd21a · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019.

TopK Language Models HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.727732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:31:39.176817Z digest=sha256:59d655242bbee7bf994c5cfaccd8785fade96ea15e238e9b0f9e5f9d2f8f08c0

Observation 757b2521-884d-4004-a072-b40518b8c712 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TopK Language Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:31:39.220992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.220992Z digest=sha256:73aeaf26ce74d29e0cd2b603fa0116ceb3902ceabb9269475913fdcf9a9f8de6

Pith citing papers

No inbound Pith citation observations are available.