Pith. sign in

Paper Citation Record · LEDGER

On Linear Representations and Pretraining Data Frequency in Language Models

As of 17 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 4 inbound Pith citation observations for arXiv:2504.12459.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12459 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:37:36.423808Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:52:44.489356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T11:06:17.562742Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9be8aa00-ec8e-4c84-a732-a1276b6a6208 · outbound

This paper cites To code or not to code? exploring impact of code in pre-training.

On Linear Representations and Pretraining Data Frequency in Language Models To code or not to code? exploring impact of code in pre-training

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.117127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.219091Z digest=sha256:cea231750bb19a2717d08dd468400e15a67a6c585881b53171fab0a440f84a6a

Observation 02a38762-7b87-431e-b172-30f41c838d6e · outbound

This paper cites Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers.

On Linear Representations and Pretraining Data Frequency in Language Models Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.222599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.222599Z digest=sha256:0539d1dbdcdb78d9df14a5af284106f9e96ebf5f4da22984a0d7785aec1a9f59

Observation 724aa831-5196-4c4f-aeb0-9e158903b4bb · outbound

This paper cites Interpreting Neural Networks through the Polytope Lens.

On Linear Representations and Pretraining Data Frequency in Language Models Interpreting Neural Networks through the Polytope Lens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.225333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.225333Z digest=sha256:71981791cf09aa881a65d176d68277f63f78d1c0fddd60c02d2fa6ce0beb2dd6

Observation 65ade467-bb03-444d-a049-8e60da7558cc · outbound

This paper cites Membership inference attacks from first principles.

On Linear Representations and Pretraining Data Frequency in Language Models Membership inference attacks from first principles

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-16T12:37:36.734597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.229182Z digest=sha256:4908fdc51f0f0e3a0fe470c0fdb383dbf6f0481ebbb5f160656e03ae78ce20b7

Observation 690dd01c-c5c6-44ff-bf86-1a0bb0f9daa4 · outbound

This paper cites Quantifying memorization across neural language models.

On Linear Representations and Pretraining Data Frequency in Language Models Quantifying memorization across neural language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.232547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.232547Z digest=sha256:64dba1183e9f931038d00a2f285a32f88abc7452d0fbc86c2cff03a8bdb4d85f

Observation abe70b6b-8a87-4ac0-9423-9ec506a8e51b · outbound

This paper cites How do large language models acquire factual knowledge during pretraining? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

On Linear Representations and Pretraining Data Frequency in Language Models How do large language models acquire factual knowledge during pretraining? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.098171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.235649Z digest=sha256:fad0a550fc8fc9262a780f5768f4a010f65191b35f746e02a8bd79e18e52f8bd

Observation 82581bb6-349b-49d6-9b5a-b1a531e61b08 · outbound

This paper cites Identifying Linear Relational Concepts in Large Language Models.

On Linear Representations and Pretraining Data Frequency in Language Models Identifying Linear Relational Concepts in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.239229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.239229Z digest=sha256:2ba9c9283e9e620b29d23119ee9b5287a4f9a37a79b218a72ff074f53588a502

Observation 5d208efe-cb70-4887-a0ca-c1b5a09a3ad8 · outbound

This paper cites Recurrent neural networks learn to store and generate sequences using non-linear representations.

On Linear Representations and Pretraining Data Frequency in Language Models Recurrent neural networks learn to store and generate sequences using non-linear representations

Reference 8

Resolution
verified exact
doi, observed 2026-08-16T12:37:36.572317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.242809Z digest=sha256:811d8b49264081d599f96bcf262e1d66c184fdaf316496da32b263da4454a542

Observation 296226ad-19b5-4b10-944c-4f34e5d2946c · outbound

This paper cites Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals.

On Linear Representations and Pretraining Data Frequency in Language Models Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-16T12:37:36.245789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.245789Z digest=sha256:e6696357f9d0849ed18b074b610eb4776d388ff8a163157eae777964436889d9

Observation c29e93ab-2b3d-42e3-9d62-b4b8eb55e9a7 · outbound

This paper cites Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions.

On Linear Representations and Pretraining Data Frequency in Language Models Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.248759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.248759Z digest=sha256:20a1556beffa8276f472d9323b60fa9b2c2bdfae674ca69484338d85bcf4755e

Observation dff16e97-9766-4084-9ddf-47eae3b4dde7 · outbound

This paper cites Smith, and Jesse Dodge.

On Linear Representations and Pretraining Data Frequency in Language Models Smith, and Jesse Dodge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.085232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.251685Z digest=sha256:bf11e1089e480185e3d89673424ed38553ee704df08e0f16a3f97abccff624c1

Observation 2128fed1-42a8-4911-864c-38476d366370 · outbound

This paper cites A mathematical framework for transformer circuits.

On Linear Representations and Pretraining Data Frequency in Language Models A mathematical framework for transformer circuits

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.255407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.255407Z digest=sha256:e200641e6768453cd2be5ffdccf41ff4b89f7095fef8f67690e9754018bfab69

Observation f9573e36-13d2-474b-9126-440c4b4dfe71 · outbound

This paper cites T - RE x: A large scale alignment of natural language with knowledge base triples.

On Linear Representations and Pretraining Data Frequency in Language Models T - RE x: A large scale alignment of natural language with knowledge base triples

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.067881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.258400Z digest=sha256:19291f73a729249362825a5f06750d048bfd0c1deb61c89758698993859da699

Observation b27cfc5b-8070-4529-8f6b-b4934ec0c0ac · outbound

This paper cites Towards Understanding Linear Word Analogies.

On Linear Representations and Pretraining Data Frequency in Language Models Towards Understanding Linear Word Analogies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.261244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.261244Z digest=sha256:42e9936b2d870e9d32f9d66063cb518f17285ca70f232c16a2cfe8e8391292e5

Observation ec363ae7-6085-4fd0-8ed9-1d9c0af324bc · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

On Linear Representations and Pretraining Data Frequency in Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.264651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.264651Z digest=sha256:71f7a32913527bea6248dad9d6808a2eacbd83b0ce47923e89ee18e6f9560fb2

Observation 16ecfce0-7101-4a7e-b91e-ef1b3bd77a7f · outbound

This paper cites Scaling and evaluating sparse autoencoders.

On Linear Representations and Pretraining Data Frequency in Language Models Scaling and evaluating sparse autoencoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.268775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.268775Z digest=sha256:2809d60ae1e85cb5c172d1c2bdf0fc01eee23ca8680310b183eb5cb92ddbaef1

Observation dacfa710-47cd-4a08-a7fc-71c7dafc86e4 · outbound

This paper cites What can transformers learn in-context? a case study of simple function classes.

On Linear Representations and Pretraining Data Frequency in Language Models What can transformers learn in-context? a case study of simple function classes

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.048571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.272216Z digest=sha256:2a2d0e502e3f9818212150eeec7c857b8f63e2b980a431ebd2fadb399e299495

Observation 8011d74e-f40c-4da3-a142-ecbac4602b57 · outbound

This paper cites Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn`t.

On Linear Representations and Pretraining Data Frequency in Language Models Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn`t

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.276369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.276369Z digest=sha256:b71aefb05c39aea10b73e0fc4875a29fc2a8ae250436e31586ac3a14da729899

Observation 0f6d728a-31b0-4b19-b183-0614e70e8c81 · outbound

This paper cites OLM o: Accelerating the science of language models.

On Linear Representations and Pretraining Data Frequency in Language Models OLM o: Accelerating the science of language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.280450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.280450Z digest=sha256:9b55b4f7cc406030832ab783bc6f42269004cd13c366fafbd2dc418da17fa5be

Observation 61de6d86-90cd-4feb-8723-69a9970c8ba9 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:37:37.036827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.284194Z digest=sha256:15bff67e1ff0221ceef2e59c9b1c28969b1046e8a71b72fd4cd9ba256cdf9fa3

Observation fc1f87f2-39a9-495b-94b9-fdc8f5cd209b · outbound

This paper cites In- Context Learning Creates Task Vectors.

On Linear Representations and Pretraining Data Frequency in Language Models In- Context Learning Creates Task Vectors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.287788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.287788Z digest=sha256:c56df1dca98623865fbb91b6b055f4d3ea093eb4bd033b4c057f2da1c099238f

Observation 935a2fb3-0694-45eb-8236-cd6caac42e6d · outbound

This paper cites Linearity of relation decoding in transformer language models.

On Linear Representations and Pretraining Data Frequency in Language Models Linearity of relation decoding in transformer language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.292061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.292061Z digest=sha256:9f75c7514e4bf0f0a32b86afae9e9f747b84c4d6fb98576cae2fabac91a48d0a

Observation ef2593e1-8db3-4dca-af3f-113b8d36f760 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

On Linear Representations and Pretraining Data Frequency in Language Models Sparse autoencoders find highly interpretable features in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.294958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.294958Z digest=sha256:783b63b2f6545e496fec681f858e2e71b6a2dbbac5b9279a580767ad778b7dbc

Observation 1820c1dd-5bb4-4953-ad68-da5d79e0ca4a · outbound

This paper cites On the origins of linear representations in large language models.

On Linear Representations and Pretraining Data Frequency in Language Models On the origins of linear representations in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.010624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.297826Z digest=sha256:820bcb53621b1056719b440a3b8469417f1dd5608bef134216f4e0c205b0867f

Observation 28864f5b-e906-4a6e-9f25-80469d09aa74 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-16T12:37:36.528778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.300566Z digest=sha256:a3cdeddc26dd253e69f883d0d0c9d9998cc428cb0d57ec06435f76ca41717c31

Observation 63c8a6f8-3f32-4c64-969c-15cb5e33744f · outbound

This paper cites Multilingual reliability and semantic structure of continuous word spaces.

On Linear Representations and Pretraining Data Frequency in Language Models Multilingual reliability and semantic structure of continuous word spaces

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.996973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.303582Z digest=sha256:e392f79227d41af6b43331d1c7306af1d8fda65b45c366e2a87d8867c709b279

Observation 9646f0d9-ee4c-458e-8ced-9bf0aaa0b3ba · outbound

This paper cites A pretrainer`s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity.

On Linear Representations and Pretraining Data Frequency in Language Models A pretrainer`s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.307341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.307341Z digest=sha256:b8731b6109d328d9196b755831933c16a2dd188514ebf44cc98443f7807b98f5

Observation 80ebb938-d5d8-4a12-b775-374d0ec56c21 · outbound

This paper cites At which training stage does code data help LLM s reasoning? In The Twelfth International Conference on Learning Representations, 2024.

On Linear Representations and Pretraining Data Frequency in Language Models At which training stage does code data help LLM s reasoning? In The Twelfth International Conference on Learning Representations, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.314919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.314919Z digest=sha256:545f1b486b382ba064c76c0e607111cf0fc8badf705800f92f06955feb1c4b93

Observation e08f2e27-05eb-4b4c-b93d-b3ca2e19131e · outbound

This paper cites When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.

On Linear Representations and Pretraining Data Frequency in Language Models When not to trust language models: Investigating effectiveness of parametric and non-parametric memories

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.317695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.317695Z digest=sha256:4dbe57967811ca7d88fd015e31b44c796bef63b861ee19982588838cc413f6b5

Observation 5148ca2b-7301-431a-9bb8-06e8e79eb9c3 · outbound

This paper cites Embers of autoregression show how large language models are shaped by the problem they are trained to solve.

On Linear Representations and Pretraining Data Frequency in Language Models Embers of autoregression show how large language models are shaped by the problem they are trained to solve

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.320670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.320670Z digest=sha256:a1f16f770a743665264138eaca84629f98e9bddced424efbeeac8a9bd5d0f6c8

Observation 7276c5e4-7105-4cd2-a26b-dbe878ecf07d · outbound

This paper cites Language models implement simple W ord2 V ec-style vector arithmetic.

On Linear Representations and Pretraining Data Frequency in Language Models Language models implement simple W ord2 V ec-style vector arithmetic

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.323521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.323521Z digest=sha256:1539f9d202ae46d0d95dae85ac876a28e666ff51443afd7dab9e281845abfad1

Observation 07fb214a-a2e4-4e3b-9bf2-a3ce395845e2 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

On Linear Representations and Pretraining Data Frequency in Language Models Efficient Estimation of Word Representations in Vector Space

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.326920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.326920Z digest=sha256:fd7b327fb39c40637ed13d75d56cffae3840db7e39eff7b931fa4da25e975e9b

Observation d423fe4a-a934-4b57-b535-2dd7941ac934 · outbound

This paper cites Distributed representations of words and phrases and their compositionality.

On Linear Representations and Pretraining Data Frequency in Language Models Distributed representations of words and phrases and their compositionality

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.970610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.330488Z digest=sha256:dfa09a8da1c4bfdd99eee7091c452100a111d259cff6ec8e419708d118239804

Observation c8898bc0-9f6a-4d58-8c0b-a8504aaba89c · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.333747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.333747Z digest=sha256:5f9cd411c5b8bf83aa231e0b95c3134ffe471416165c3489652a672a783c0aa9

Observation 002554d9-8b63-4da5-9dec-f759a6602b8c · outbound

This paper cites Zoom in: An introduction to circuits.

On Linear Representations and Pretraining Data Frequency in Language Models Zoom in: An introduction to circuits

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.959877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.336933Z digest=sha256:ca3350a207159562a707f8b7c0bfe60ff5ca9972b447e7892e00e7d77312c650

Observation f021bf14-4dde-4a74-9088-6eb57540e730 · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

On Linear Representations and Pretraining Data Frequency in Language Models Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.341261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.341261Z digest=sha256:0a24d730d6ba1943ea189a026a8ef25a72598f7ed38aead0355dd13c0f41a5fa

Observation c529b3bc-2e8f-452b-bdf4-4a5c14891550 · outbound

This paper cites Learning Hierarchical Structures with Linear Relational Embedding.

On Linear Representations and Pretraining Data Frequency in Language Models Learning Hierarchical Structures with Linear Relational Embedding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.940678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.344829Z digest=sha256:4e04e7585f02c7ca1b0a7f60296966ec17bf745b36de7116bc463b8cb46254ac

Observation 059f7a5b-ed0e-4560-abea-b6611f4a5ae9 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

On Linear Representations and Pretraining Data Frequency in Language Models The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.929566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.349063Z digest=sha256:0e59cc4c30bd5536c66d09cbaf27b3f191ac8f46012425d73151b2078ff6649c

Observation 78bf24e7-b5fc-47f4-98f7-6dbc93b669ab · outbound

This paper cites G lo V e: Global vectors for word representation.

On Linear Representations and Pretraining Data Frequency in Language Models G lo V e: Global vectors for word representation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.352815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.352815Z digest=sha256:bf9998f58548b2a921cf868cf8d52257ac4d956cc9b0a597d79fa1ed4f74d240

Observation e8e3a177-e391-4a5d-b57d-a290d40dad70 · outbound

This paper cites Null it out: Guarding protected attributes by iterative nullspace projection.

On Linear Representations and Pretraining Data Frequency in Language Models Null it out: Guarding protected attributes by iterative nullspace projection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.917292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.356294Z digest=sha256:3cbbd1d098cc19c2f8c8a282b6ca324e0b4061afe519c8d82935c38d192222b9

Observation a78393f9-dd28-4252-89bc-4b33ca76a95d · outbound

This paper cites Impact of pretraining term frequencies on few-shot numerical reasoning.

On Linear Representations and Pretraining Data Frequency in Language Models Impact of pretraining term frequencies on few-shot numerical reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.359724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.359724Z digest=sha256:55b370ee69bbb4d1c2e756500cee8b2247ff2a968b2ee75fc56b12466e7cf83d

Observation 77363836-c1b2-4098-b3eb-386f16046404 · outbound

This paper cites Backtracking mathematical reasoning of language models to the pretraining data.

On Linear Representations and Pretraining Data Frequency in Language Models Backtracking mathematical reasoning of language models to the pretraining data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.905434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.362674Z digest=sha256:d8207611255957614ca5603bcfabc0e8728ab738f24c4b30bababb5495976fcf

Observation 439fd0fe-b768-4db5-8271-0a4eea112db9 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

On Linear Representations and Pretraining Data Frequency in Language Models Steering llama 2 via contrastive activation addition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.365962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.365962Z digest=sha256:9df81bbd4a4b03649d15b391417ea72688a7120c4b445086bd095ba986f5cd3e

Observation f2046014-23a3-4e24-a48e-82f066f28c8c · outbound

This paper cites Salton, A.

On Linear Representations and Pretraining Data Frequency in Language Models Salton, A

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.369689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.369689Z digest=sha256:2591a53a238b6078f0c4a699c348f7c18042d80b0e7204c4fa8d483a6ac06156

Observation dd67bddf-8ca1-4c1e-b87e-44f74ebb32a4 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.372202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.372202Z digest=sha256:794bf096bc14b34df373aa8ee8e3cf852710b4c1e1563263b4b9d7cfeac2cc60

Observation 3f630405-b32b-49dd-a1de-c818931f275e · outbound

This paper cites The bias amplification paradox in text-to-image generation.

On Linear Representations and Pretraining Data Frequency in Language Models The bias amplification paradox in text-to-image generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.375367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.375367Z digest=sha256:91d39d64d21fc3c988446904d113d17b680526a9b1b80656e48b6b3a61fdaa66

Observation f37685d7-e8c3-4065-99d2-e447deaff3da · outbound

This paper cites Detecting pretraining data from large language models.

On Linear Representations and Pretraining Data Frequency in Language Models Detecting pretraining data from large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.378430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.378430Z digest=sha256:f33b98afd4bb36dfd570c1cb94760850a2820a60d20557374f94bbbe02763dd1

Observation 707bccd1-206e-49a4-ad5f-103f29e67cc9 · outbound

This paper cites Membership inference attacks against machine learning models.

On Linear Representations and Pretraining Data Frequency in Language Models Membership inference attacks against machine learning models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.381493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.381493Z digest=sha256:0b79b3d8457f1cf8e9eabc3a746f612ea851b5aa5001d3c93ed21a6ec36fd104

Observation 8bdb9137-e7d0-4ed3-b297-acb6ea1f6d2a · outbound

This paper cites The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models.

On Linear Representations and Pretraining Data Frequency in Language Models The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.385139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.385139Z digest=sha256:db48c9999c89874f249674e4b4edb37cccbc46dee5cbf08252f8eca27270d23f

Observation 87c6a1e0-ac1d-4729-875f-966a3e12ef57 · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

On Linear Representations and Pretraining Data Frequency in Language Models Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.388267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.388267Z digest=sha256:645b50b85603251ec03f4f230012157811ebb5e3261f21cae3319ba874ed2cf0

Observation c8182546-7004-4d5e-b677-b7a59e917a92 · outbound

This paper cites Extracting Latent Steering Vectors from Pretrained Language Models.

On Linear Representations and Pretraining Data Frequency in Language Models Extracting Latent Steering Vectors from Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.390855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.390855Z digest=sha256:d08946f370629e7f5f93d2dbde5a0388c7bef9f42943b0d1ce8b4aaa0e8be46a

Observation 3bf08ac5-ee0c-47ea-af86-ba018a230d97 · outbound

This paper cites Formalizing and Estimating Distribution Inference Risks.

On Linear Representations and Pretraining Data Frequency in Language Models Formalizing and Estimating Distribution Inference Risks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.393954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.393954Z digest=sha256:638143a7de4c4d4beca276b72732f0ba9c48ae7ecf8b13cb4055d9a7c2a94b78

Observation ab1dba8d-4911-47ac-b555-1fe3ce592346 · outbound

This paper cites Scaling Monosemanticity : Extracting Interpretable Features from Claude 3 Sonnet.

On Linear Representations and Pretraining Data Frequency in Language Models Scaling Monosemanticity : Extracting Interpretable Features from Claude 3 Sonnet

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.880815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.397117Z digest=sha256:8784b2cd92960a2ad63408525d48a9d857c9b53fe4bb5fa2019738979e97ab3a

Observation c10e049a-c1a9-429b-a62f-f7d6d78e05fd · outbound

This paper cites Function vectors in large language models.

On Linear Representations and Pretraining Data Frequency in Language Models Function vectors in large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.399886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.399886Z digest=sha256:4691225693d06ceda1cb9d8c23f0c8831c4be4e03fa704f963a90cb80765f1c6

Observation 980dfd8b-4b2f-4b44-866e-321f219ce919 · outbound

This paper cites GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model.

On Linear Representations and Pretraining Data Frequency in Language Models GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.402311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.402311Z digest=sha256:8960ad67e6a4a2a90a67990a7163707b7100ef1dffa8e9456b22a88190720129

Observation 2d105eac-e5ca-42f4-b4de-233b54968e97 · outbound

This paper cites Understanding reasoning ability of language models from the perspective of reasoning paths aggregation.

On Linear Representations and Pretraining Data Frequency in Language Models Understanding reasoning ability of language models from the perspective of reasoning paths aggregation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.855643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.405447Z digest=sha256:f53bdf6fd9b52f69217427d6ec9e4a0c8948092b63461185e53e232becefac1a

Observation a2631398-a306-4511-9fc1-44dc66fb242d · outbound

This paper cites Generalization v.s.

On Linear Representations and Pretraining Data Frequency in Language Models Generalization v.s

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.408579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.408579Z digest=sha256:a8e32715922369a941c19c2b03f931a86f522da2c30789eef1052e4adcf7091e

Observation 2f03f27a-ce68-4dbd-8a02-dd110def5698 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.

On Linear Representations and Pretraining Data Frequency in Language Models Doremi: Optimizing data mixtures speeds up language model pretraining

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.411183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.411183Z digest=sha256:470edd60b883398872d6f59f0740b5091f835a3559168a08480b1f0037230631

Observation c1b8bedb-0c0b-4e4e-aca8-801ac0042f62 · outbound

This paper cites write newline.

On Linear Representations and Pretraining Data Frequency in Language Models write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.413796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.413796Z digest=sha256:0abc478eb872b7cb02e577391e11c82a44896120d97d408887775c4c24d54c6d

Observation 10cc502b-2936-448a-a369-1dcf261fe555 · outbound

This paper cites @esa (Ref.

On Linear Representations and Pretraining Data Frequency in Language Models @esa (Ref

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.417232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.417232Z digest=sha256:2240642fe9b1f79c3e9854696668fc5f004f4165f1356e4ea10fe0fd9e68d6a5

Observation 9f53c3e7-0937-4ab4-9c36-cbe7f4006e84 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.420669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.420669Z digest=sha256:4a548a6404be177369b34d18cba021f4520380c60209e7cfe9408a7c8f0efe1b

Observation 4fd38eb4-93c0-4589-96fe-51cfee1bf51e · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.423808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.423808Z digest=sha256:f0a6ed9518d9497be8231c7a2c78b24e554907351e8f972d639cd7d44c940633

Pith citing papers

Observation eb9371a3-71e5-430c-a26b-c3960bc65be3 · inbound

How Do Language Models Compose Functions? cites this paper.

How Do Language Models Compose Functions? On Linear Representations and Pretraining Data Frequency in Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.566205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T11:04:09.179536Z digest=sha256:d2327fd94d29c91dbba55aaaaac1a3624a2f3f20f5e8c30f1bb53064b07c478d

Observation b4298fba-acae-4c14-886a-9dc5eb2345b5 · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space On Linear Representations and Pretraining Data Frequency in Language Models

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:19.242265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:814b37fa07fb6a90da2345eec44f9feaae4998cccc071e376231a0f7397e323f

Observation 72086d85-fd50-4d02-ae8f-6cd0aac35639 · inbound

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering cites this paper.

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering On Linear Representations and Pretraining Data Frequency in Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T04:52:44.489356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:52:44.489356Z digest=sha256:a0c1cc3094e103afc3c47299118fd03b5a40c65e642397b1b6bdd2759c35fd41

Observation e530c710-5451-478f-b005-66aa643c563f · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining On Linear Representations and Pretraining Data Frequency in Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:05.090555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:05.090555Z digest=sha256:3bc92fef5c9ec1cd30c0fdb8043f9cb68b55594f38d8338cff0db0cb416176dd