Pith. sign in

Paper Citation Record · LEDGER

TopK Language Models

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21468 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:31:39.220992Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 789d64b4-69c1-41fb-80dd-51772945f630 · outbound

This paper cites How can we be so dense? the benefits of using highly sparse representations, 2019.

TopK Language Models How can we be so dense? the benefits of using highly sparse representations, 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.922624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:19.862785Z digest=sha256:408d35c46a1c07de5034b94158234f66d82f96d0197aaebdab2d16b6b4d8e86c

Observation 252a90f8-e166-4ccf-8aa8-0427bd7175dd · outbound

This paper cites Generating long sequences with sparse transformers, 2019.

TopK Language Models Generating long sequences with sparse transformers, 2019

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.787577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.787577Z digest=sha256:6d41960b22f14c0953f2fdcc9f2bf56094089c58c87f341cd48cecfa6f1f50ea

Observation 674b7941-d134-4ec2-b182-617c4b245cb5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

TopK Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.894657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.894657Z digest=sha256:08752f98c3a2220cdfc5601b17e27d35655ee82477bb520b9bc3554419b68fcd

Observation 7a6bcd59-399b-4e81-a450-405ea058ef91 · outbound

This paper cites [Full Post] Progress Update #1 from the GDM Mech Interp Team.

TopK Language Models [Full Post] Progress Update #1 from the GDM Mech Interp Team

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.790819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.076386Z digest=sha256:54f9853a8e5c5a0d93fa134b252d440a74755c0244c40a3897755b2b6238632c

Observation ba9b874e-ce44-4514-9642-ba67d7764f98 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models, 2023.

TopK Language Models Sparse autoencoders find highly interpretable features in language models, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.158355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.158355Z digest=sha256:79648a9b04d368a7d76f2d72e5814e7a7ecfaf625ec081a159d3e1837551a2bb

Observation f0204e56-9a10-4e7c-947b-1d805f936fb0 · outbound

This paper cites Rigging the lottery: Making all tickets winners.

TopK Language Models Rigging the lottery: Making all tickets winners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.685373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.321207Z digest=sha256:2d102dcbeba19cdff564418a6de71cf3d8c003ca41b8784496200bedb978840a

Observation 30ea9f6d-00f7-4f49-964f-0f3dc6a16879 · outbound

This paper cites Sparsity in transformers: A systematic literature review.

TopK Language Models Sparsity in transformers: A systematic literature review

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.560982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.434488Z digest=sha256:bec5c39440b3552dbc500766652fb6f4f291583cfbcd2c3cfe3a03ff9a39b9a8

Observation 72d4a652-94cd-4b34-a2aa-d89f288ff9b0 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

TopK Language Models Detecting hallucinations in large language models using semantic entropy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.530869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.530869Z digest=sha256:2963621f0042cd5693eaed393ec82e17eca31fd6c77158fe40f6ea4b7879cae3

Observation 6f137638-06ec-406b-9ce0-061a3bf57b53 · outbound

This paper cites Scaling and evaluating sparse autoencoders, 2024.

TopK Language Models Scaling and evaluating sparse autoencoders, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.452327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.680217Z digest=sha256:01511cc29b961ce41d852b849f4eda4068e03dac7a74789b67e2e99b89cbb509

Observation f4e0c28a-7209-4240-955b-333c6e0ea7da · outbound

This paper cites The language model evaluation harness, 2024.

TopK Language Models The language model evaluation harness, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.370392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.732571Z digest=sha256:9c63d2f1733d0a6d9d5bc96c8759e97d7f00e60830714e0fe0e7b91b1bc7a0ac

Observation ee4dd635-db57-410b-9324-1787e68f3a5c · outbound

This paper cites Causal Abstractions of Neural Networks.

TopK Language Models Causal Abstractions of Neural Networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.255996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.794053Z digest=sha256:d28c6a6ae3fc77c7045533513bd7d6bd096c0c9ab7b117dddb45e7b876d791b4

Observation 04f5fb92-8a51-49c8-b8d4-d99b82567568 · outbound

This paper cites The State of Sparse Training in Deep Reinforcement Learning.

TopK Language Models The State of Sparse Training in Deep Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.158063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.872432Z digest=sha256:95f2a3a62a2d8d290d38e59eed90b68d1e308de831813a53a843121685f83499

Observation aa35ef23-a1e2-4365-8576-b6daf9a5c923 · outbound

This paper cites Memory-efficient transformers via top-k attention, 2021.

TopK Language Models Memory-efficient transformers via top-k attention, 2021

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.047952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.032941Z digest=sha256:91b86951e991c47faa95c511ad4d338dfb755e9b5d894026984e1082dfd7ceca

Observation 3f7984cd-049b-4bda-86f3-9384fd31b028 · outbound

This paper cites Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024.

TopK Language Models Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.933604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.096247Z digest=sha256:8fa9ca53d910b98c2cf96c10efdca29886396746b8422e55337df29fd7e3c9f2

Observation cdb4a400-794b-422a-b9db-23e9e5554868 · outbound

This paper cites Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba.

TopK Language Models Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.826679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.144498Z digest=sha256:12b05cdd0f457563855d85edfebc73290a2f65e8a1bb72553d7e606859928dd9

Observation 93a23817-5ae8-4063-9a61-cc7a57c7407d · outbound

This paper cites Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks.

TopK Language Models Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:39.565510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.250242Z digest=sha256:cf9115dc74cb9e1b5058c92bfd384ed93a742e554d5c09384454cc08d1d4c2af

Observation 431d5c70-f9cd-434d-b265-d0c8d173ceb6 · outbound

This paper cites How llms learn: Tracing internal representations with sparse autoencoders, 2025.

TopK Language Models How llms learn: Tracing internal representations with sparse autoencoders, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.727068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.341495Z digest=sha256:ee86715945c0ee666dc218ebe08fce65cc517cff5451bf0ca97c160df38ccd91

Observation d3c0bed0-dcef-4bad-a2a3-7229280fc0ad · outbound

This paper cites Sparse is enough in scaling transformers, 2021.

TopK Language Models Sparse is enough in scaling transformers, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.615792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.403525Z digest=sha256:9d9d64d899a979754358d169c1076d9a8bbd4be8c091aa799604cf44e796692d

Observation d7578cb5-abb8-45ff-b7d9-85726b9abc70 · outbound

This paper cites Top-KAST: Top-K Always Sparse Training.

TopK Language Models Top-KAST: Top-K Always Sparse Training

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:31:39.411734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.509531Z digest=sha256:f73e8753f22dc963792a567a4c8ff82f436bd45bd04c080a67dc14c3284dd666

Observation 4f91072b-f53e-4f6f-b9bf-d3d8d7b64d13 · outbound

This paper cites Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025.

TopK Language Models Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.505754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.629380Z digest=sha256:8e16b2af798b718b0370afbd141c589354ffb55eb72f1669a55f33058bc98964

Observation 1d8d7c24-9783-465d-9c80-be28083f430a · outbound

This paper cites Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025.

TopK Language Models Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.327447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.737124Z digest=sha256:de561557973f0cf38f690ddee69252fba950b048f1b6dbc61e38505cfbec7085

Observation 88cbb3b3-8ac3-4ba0-8158-3b1251dba858 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

TopK Language Models Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.842439Z digest=sha256:8f2771ed45d6fb18f1875a6a78aa9bf097b49609e615171f229217d750a95cc4

Observation 3edfc9c7-94f6-4281-a590-880224887ad8 · outbound

This paper cites Soft Threshold Weight Reparameterization for Learnable Sparsity.

TopK Language Models Soft Threshold Weight Reparameterization for Learnable Sparsity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.002177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.925606Z digest=sha256:da5704099a14292c55c81ab2e44e97d4c1a9515c1be3d2ffb8fc82e573c257d7

Observation a05d54fd-ae7b-4413-b870-4cecc3413d54 · outbound

This paper cites Sparse autoencoders do not find canonical units of analysis, 2025.

TopK Language Models Sparse autoencoders do not find canonical units of analysis, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.807089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.983433Z digest=sha256:20a9de1ab9b38fea580392c984fb9f6ae44a51ab86bfc3a88cc9b6d3a5c2eb92

Observation 5f4fba51-bb4a-4ebd-9a82-f17b303fec40 · outbound

This paper cites Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024.

TopK Language Models Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.056013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.056013Z digest=sha256:f56fd76e3b20e7eea260162dc284526a2deac6561fb218f6128ddf5f30003b20

Observation 0e04cee2-46f0-46d1-ade5-282892edb508 · outbound

This paper cites Decoupled Weight Decay Regularization.

TopK Language Models Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.178870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.178870Z digest=sha256:67385f18a5c8bec041b712323f59784633630a7adf98a3ddc93170222140559e

Observation dfa20b76-4411-44f4-807f-6be32170786e · outbound

This paper cites Learning sparse neural networks through l_0 regularization.

TopK Language Models Learning sparse neural networks through l_0 regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.623753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.243550Z digest=sha256:caf3bece0c9feed96ff28c8bf84d2ece5a0d8fea06312583e96bbbbfd2659212

Observation 8a792f02-b2c6-45e8-b471-bb1f4f042e41 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

TopK Language Models Fineweb-edu: the finest collection of educational content, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.336793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.336793Z digest=sha256:0089a34028976afeb64addb68de8cf71ae0e9699ff1eda6df8eb11d8f99a9bce

Observation c81c18f1-a5e1-419b-878a-9a4b2a29115b · outbound

This paper cites Winner-take-all autoencoders.

TopK Language Models Winner-take-all autoencoders

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.399846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.409944Z digest=sha256:a6c2cdf27b192f5fca6a806d5a4f76508d76ab787c0a1c063e5d465a15be4f61

Observation cf1a13db-df4e-4f85-8d87-3a1f51a45d56 · outbound

This paper cites Pointer Sentinel Mixture Models.

TopK Language Models Pointer Sentinel Mixture Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.252639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.516145Z digest=sha256:fcb352c2b387c07c5a5af31ac89c75c7db54525c9f8dd54317a0acf6df0a2b0f

Observation 6cc3d4c0-c29f-45dc-b26e-406dba71cc93 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

TopK Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.067358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.572894Z digest=sha256:791566f8a5d47946d7e3307f4c000dde504134c96e3fe17d108e89f21dcd8321

Observation 69e9c625-3b6b-45ee-883a-cc68e008da50 · outbound

This paper cites Nguyen, Madeleine Gibescu, Antonio Liotta, and et al.

TopK Language Models Nguyen, Madeleine Gibescu, Antonio Liotta, and et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.886349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.620117Z digest=sha256:5cd4432e1f426ba021c2140d19dd54155361e4b52359ae295252a02038ce6bfe

Observation 2a719d09-833e-4914-8ee3-0b74cd9203b5 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

TopK Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.688817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.666381Z digest=sha256:36814ee06e5a5d3c6e1f65685137b88fd37c7bb8ce32aa6629d61d13c77e28f2

Observation 6dd0f983-9ab7-4e9c-8faf-5f2e127b6e35 · outbound

This paper cites Sparse autoencoders trained on the same data learn different features, 2025.

TopK Language Models Sparse autoencoders trained on the same data learn different features, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.469546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.703915Z digest=sha256:11b613fd0a18b83a8028a07de6f11df1623b9b74472d71f9ed7a0a19a8e65088

Observation 0754bfa7-a587-49c0-ab97-1f49c4ce81b1 · outbound

This paper cites Automatically interpreting millions of features in large language models, 2024.

TopK Language Models Automatically interpreting millions of features in large language models, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.268686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.798385Z digest=sha256:d63dbf95d74886c5f164d05e7362105f7da52992b47f05b4cf600d2a304103a7

Observation a1125760-f628-4ca2-9ebb-4d6f608633bf · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

TopK Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.035447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.854065Z digest=sha256:c76867e2cb53e92e611d3c1440df6475976845468d98062695e65dd8e6fe1d7b

Observation d02c0341-8289-4f6a-9e93-7c948279a759 · outbound

This paper cites Improving dictionary learning with gated sparse autoencoders, 2024.

TopK Language Models Improving dictionary learning with gated sparse autoencoders, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.972247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.972247Z digest=sha256:f3beea778390ff6b6dda6685a4faf12b161104ce3bc62f15fd90597da824b433

Observation acedd645-2f42-434a-8ef1-aebff0af90bb · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale.

TopK Language Models Winogrande: an adversarial winograd schema challenge at scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.142907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.142907Z digest=sha256:6a2d79683718ebad5cfa9ad7d2726611454c66438c8091ee30d01de1f4f94e88

Observation 8f6dab73-abb4-4500-b8d5-f326759064a8 · outbound

This paper cites Taking features out of superposition with sparse autoencoders.

TopK Language Models Taking features out of superposition with sparse autoencoders

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.851527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.258973Z digest=sha256:fbaacb9749cf81cab0ec9f68a6115e7952f64c6e272176a4d67f167030b6b75f

Observation ed8b5f35-897e-4c8d-9f38-e00bfacbbac6 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017.

TopK Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.370614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.370614Z digest=sha256:bf13379607a7ae5c6c3d61e1c4b4338b8e2271fad49bdfffb22b9211afae5485

Observation 1ddad81e-ea7a-41e9-b420-c9e5d64197bd · outbound

This paper cites A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025.

TopK Language Models A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.670077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.509917Z digest=sha256:0850306bb9d87f5b258778c141e21005300b563edcd50896a96d0cd88292c525

Observation 9c0496bf-71fc-42ea-a76c-b215fc6f7781 · outbound

This paper cites Codebook Features: Sparse and Discrete Interpretability for Neural Networks.

TopK Language Models Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.680344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.680344Z digest=sha256:1efa5c7e556e8310b40363961019f9601820cef23baca98c102dd74042ab318d

Observation e1b8015f-8dca-48ee-a6d5-94e0197c7317 · outbound

This paper cites Daniel Freeman, Theodore R.

TopK Language Models Daniel Freeman, Theodore R

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.462326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.794923Z digest=sha256:8bf49354a2d2cd348b21fd44da87b2ac47c7e83cbae510e430e27a42a0f2db3e

Observation 4e902680-9768-47c0-aae7-a1a9b5f98c7c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TopK Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.890079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.890079Z digest=sha256:0c383dbaae2f33765f0f74b788a3035e23ff18f02c117186704f1e63be9a7657

Observation 3ea79e63-1c97-4e4a-ad6c-80dc996c16ab · outbound

This paper cites Meta Lingua: A minimal PyTorch LLM training library, 2024.

TopK Language Models Meta Lingua: A minimal PyTorch LLM training library, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.282755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.998159Z digest=sha256:974a1782d750a457f66d467db555b80b24da3d0261e6aeca846e94df3023d243

Observation 373449a0-f1a3-45e0-b63a-24c6155ab499 · outbound

This paper cites Tracking the feature dynamics in llm training: A mechanistic study, 2025.

TopK Language Models Tracking the feature dynamics in llm training: A mechanistic study, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.093824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:39.060459Z digest=sha256:204c3a058c8a5b622f8532185008bc53c1059db0b6a6760a0c9e057d90e83059

Observation b26d6674-e1c9-49fd-a3d2-63892221c21a · outbound

This paper cites STEP: Staged parameter-efficient pre-training for large language models.

TopK Language Models STEP: Staged parameter-efficient pre-training for large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.921354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:39.121952Z digest=sha256:50c2267ef855822323c372d3faff491c84c27509447e648c09bbc927f0042a50

Observation 4a13895a-5f6b-42e3-a2c2-6056c77dd21a · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019.

TopK Language Models HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.727732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:39.176817Z digest=sha256:c798b6da8cf79a47e1b32913bfa6b141d20bb1585d93de44eceb6e014967d2ca

Observation 757b2521-884d-4004-a072-b40518b8c712 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TopK Language Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:31:39.220992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.220992Z digest=sha256:0f83bbd6ce7caf6337edabc37b1d1efc0517e8055dff309462611e98e266c99c

Pith citing papers

No inbound Pith citation observations are available.