Pith. sign in

Paper Citation Record · LEDGER

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2505.07293.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07293 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:23:36.717737Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T12:57:26.458109Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T06:07:41.034120Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved49
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15052427-48b9-41a9-8b25-c77efce244c3 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.342892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.342892Z digest=sha256:003add1d7e942a0310014be5b038d5ccd3f9b10be1ddd6faac05f44c96c8f414

Observation 569a43c7-96b9-4512-8f85-cdd643740b12 · outbound

This paper cites Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.349667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.349667Z digest=sha256:ac74e6fce188cb07435d189b1338a17318a5a8de22f2dbba828cca548c2275fb

Observation b7ade84b-4b54-4018-a42e-bee1c430bb16 · outbound

This paper cites Smollm-corpus,.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Smollm-corpus,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.355299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.355299Z digest=sha256:873a9a94b014da02980782a46b214f36b5972fe3b6dcf45ae84156cbbead27d6

Observation 2c76a108-f35f-40b1-903f-9d77feff11d3 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.368664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.368664Z digest=sha256:369cf13791f0ed0a98ee6b005a311507a2fb7ba583c5d611276a6ede074b0279

Observation aed61845-245e-47aa-98eb-f335b73c0347 · outbound

This paper cites Towards monoseman- ticity: Decomposing language models with dictionary learning.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Towards monoseman- ticity: Decomposing language models with dictionary learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.690423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:35.376604Z digest=sha256:4d2f07044e2357d5cd465e638494c6158f06824a2ea3050cb3fb1fc5f09f5076

Observation 3cf91f5d-d6c8-42ce-925d-e88747df4240 · outbound

This paper cites Efficient Intent Detection with Dual Sentence Encoders.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Efficient Intent Detection with Dual Sentence Encoders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.382927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.382927Z digest=sha256:ff6d9f6b071a3772aed6fe6eafe95b1e2b3a0b60a62779a6fae9f279298c3650

Observation 1baea1ab-b71e-4cca-8a05-cd76f6d155e2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.471433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.471433Z digest=sha256:bcc06a786d42c09e6279f25fb821e8fc0682db9a37f442cf803fdc68e8f3cb26

Observation c6c97e77-ca05-401c-85fe-9ce465d75f56 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.565205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.565205Z digest=sha256:52e7d0c8533de18dd14fd9acf542bb90d38440efd018a4aee05abbb091e9671e

Observation c64a20fe-1d9e-4398-9eef-1b2ae8502e15 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.679161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.679161Z digest=sha256:12a8f1f1e9e179746538d2260f2b8819ffbe002ed0f1d5645ce013aba46e3403

Observation 4a00cb4c-a9d6-484b-9a20-5c363045ed9d · outbound

This paper cites DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.686363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.686363Z digest=sha256:03d38e0752e74041c3776e671230d671989d0ebd716dfb124d926e5daee8ea54

Observation b36fbfd1-88ab-4571-8ea4-5434f4336146 · outbound

This paper cites Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.691752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.691752Z digest=sha256:2afc03f5217b9ad47c3b065e040d32806928b43cb9ffd0f73c031ccd78ce37d2

Observation 6a40318b-0c9f-4437-9189-f6318105ea54 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Transformer Feed-Forward Layers Are Key-Value Memories

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.696873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.696873Z digest=sha256:53d6807cd2bbe7f6346d265ea6c46fdf06feaaed3fbcac00b04f5af81701b231

Observation 7acfe777-81cb-4bb6-a786-6e249ad97c33 · outbound

This paper cites The Llama 3 Herd of Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.701654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.701654Z digest=sha256:3c6b9f0d2016913b5507460e989ace7df84288de04d83db168433f73dea0f427

Observation a88aba2a-ac26-4803-9af0-f023b5bd3126 · outbound

This paper cites Optimizing Pretraining Data Mixtures with LLM-Estimated Utility.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Optimizing Pretraining Data Mixtures with LLM-Estimated Utility

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.706278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.706278Z digest=sha256:b38c1a7285311e31d0e97336bfb02e32938363fee18e7e992bf4597e63b560bd

Observation 536a8b14-8458-4a31-87c7-0ef0a3cb05b5 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Measuring Massive Multitask Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.755937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.755937Z digest=sha256:f7062c801112dd0f74cbe79ad270bed41c2ec802010a1812fb8ecd4258d9da8b

Observation 2d03563a-249c-4f5d-b769-cf0bf78000ac · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.865394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.865394Z digest=sha256:1326c6922ea026795b84f1f17cc9c62d7c448d77b8203d6837d0f97884c75a86

Observation bdc98545-8281-4dcb-b2a6-0c9efa88a7f3 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Distilling the Knowledge in a Neural Network

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.972358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.972358Z digest=sha256:4713665d53bdb54300d55b687a0d0b8dab653349eb1ca7c6673daedf7fe7b09b

Observation 5cbba9c2-fc87-44cc-907f-07727ec3d6d6 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.976986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.976986Z digest=sha256:de9b4e4fb26319a0805f928a47466e2fa33a6472c7e762d3e4f33086e8513fd0

Observation 1306cc60-deb5-40fb-b1bd-9a2b326454ae · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.981579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.981579Z digest=sha256:338d805f6e1af7a3fe968c28ddb059572e55dc23e70c963c4e9c709018101f55

Observation 7624a982-d77e-4261-b3ba-4d94b28278ca · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.985894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.985894Z digest=sha256:efa4c678d21f7701439353968431e90357440bb00310370cebedea955f178c19

Observation d18e8cd2-ecdc-4426-a0e5-248831889a75 · outbound

This paper cites FastText.zip: Compressing text classification models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection FastText.zip: Compressing text classification models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.991918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.991918Z digest=sha256:a833342a21a7fe15f9bb67baf469b42397deff201923ad58e09c0b88af0b4a0e

Observation 9bce7833-f631-49c4-98fb-4c6689a22566 · outbound

This paper cites The mirrored influence hypothesis: Efficient data influence estimation by harnessing forward passes.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection The mirrored influence hypothesis: Efficient data influence estimation by harnessing forward passes

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.667541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:35.996508Z digest=sha256:dbd1c483f0e89207739e2755a274f900d7692dd0e02ffaeae842adb6d8a2b7f7

Observation 305aafe7-e1ab-470a-8712-979f547e8a11 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.000905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.000905Z digest=sha256:d5d97a6bb14af23724803a64b054f537ec0b9ba147469a268adfd8ca33a73b8e

Observation 5dbcc96b-45f7-4856-9eec-336deb053f64 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Datacomp-lm: In search of the next generation of training sets for language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.004896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.004896Z digest=sha256:8a13266b291530aff63342cff1ac158949c3f3bbb194304afee6e32bb2be12c5

Observation 4bc5fe5b-f274-4d54-b097-8cebf7207d18 · outbound

This paper cites ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.012927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.012927Z digest=sha256:2e1c854a02f98773255694da9fe3df4e4039464992066f8483d96f8627591947

Observation bf3bc0d5-060e-4834-b6ee-e68a1b4bf6a3 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Rho-1: Not All Tokens Are What You Need

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.017738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.017738Z digest=sha256:da134b46466c6e54209b38fbca397cd7eb7e85ce2384959f881b0efb6e8558db

Observation 163f28b4-f51b-4165-a90c-3ea13af6f64c · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.073298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.073298Z digest=sha256:f14f8648a462147184bb85d23a5d2a12fde768352898da6a6aa7af1686b2269f

Observation de5d1edb-5f33-423e-9b9f-477cb03c0a84 · outbound

This paper cites Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.181667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.181667Z digest=sha256:62ba43045b23fdde2024ffee75c111bfc7b70991c17b150afa39c26cbf34e1e7

Observation d43b5d7a-c0cd-4cb6-b09f-f7d93bb999d9 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.303495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.303495Z digest=sha256:4c3f3dcb5aae2c45f6ea7fcfd5c74be1a42c757626fc0c16d18697f5964f47cf

Observation d5fc0da0-896e-4e02-8ebf-acdffd5c36c4 · outbound

This paper cites 2 OLMo 2 Furious.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection 2 OLMo 2 Furious

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.309882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.309882Z digest=sha256:31b66988cee814bddb2d0a0964ceef0dd9fbf8b72d0bf791bc061be5ccee5159

Observation d7c0278b-4f75-47d2-bd93-13c13b5b2a1a · outbound

This paper cites In-context Learning and Induction Heads.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection In-context Learning and Induction Heads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.314555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.314555Z digest=sha256:9be56867178cdf467034764df5e56b57458d398484f6b3c1e662a2cae142b039

Observation 549d502c-fe53-4c79-8ba9-790de5786f38 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.Advancesin Neural Information Processing Systems, 37:30811–30849, 2024.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection The fineweb datasets: Decanting the web for the finest text data at scale.Advancesin Neural Information Processing Systems, 37:30811–30849, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.318551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.318551Z digest=sha256:362f3760db9596e89ed5da07559c469e3734f9b9098b90637c174483a40e770c

Observation 8e477ee4-2df2-4ae2-93ce-7c1326bd5a14 · outbound

This paper cites DataMan: Data Manager for Pre-training Large Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection DataMan: Data Manager for Pre-training Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.323015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.323015Z digest=sha256:6e0f95fa7216438ac948b7b69e872f2f77c3280a1a104d3deabe8c34d63d855e

Observation eccd86e7-1eeb-4d3e-a5d7-efc6425c456b · outbound

This paper cites CLongEval: A Chinese Benchmark for Evaluating Long-Context Large Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection CLongEval: A Chinese Benchmark for Evaluating Long-Context Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:23:37.109571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:36.326984Z digest=sha256:9c9426d09fdae1568ccf56eecdd09c77d19273c8ce88754419018231e5c7d0aa

Observation d9e3d4bc-8ec3-44a2-81f5-a5fe463174d9 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.331205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.331205Z digest=sha256:68199e762fb5f716589bfaedcfe9d834232009d2944a4dbcee166be25850a059

Observation b1a710f5-84fa-4872-b11f-228ddd6ecf9a · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.335197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.335197Z digest=sha256:224b63462282bfc2a22ff368e796ad9d9d13ae96eb61230990eab6afde013408

Observation e4f4725c-4c7c-4f97-a2d8-ab229eb7892a · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.433043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.433043Z digest=sha256:c6a9a6f7b3c39f59bced20d40061ee615095c5f9d42fac9401dbf08023da6000

Observation d00461b2-4989-4b1b-bdf6-2e1ae719bfb0 · outbound

This paper cites Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.485660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.485660Z digest=sha256:326de08f34d1e103c75ba421c54418d719492a2459f5e58e01a7b45cbd1e5e69

Observation 1215aaac-4870-451d-8d1f-5006b754c1fa · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.539919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.539919Z digest=sha256:806527e80bed8121bc2545dd11ff6eebef23de4db73873ed681ad8a12ced799c

Observation 27bba49a-c6ab-4884-8771-3e1f26187e7e · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.545689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.545689Z digest=sha256:be459e4ec7c77f2fef552382f40203948732b107512db0471e146a525387bee7

Observation 0e8b2492-d185-4735-80b9-9012088498c6 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.551402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.551402Z digest=sha256:4b833f09ff92d7aa6f704411929cf9bbd3db4996be869075641a7a333a909e87

Observation 27ee5e9c-59fb-450c-ae64-9de5435ad704 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.556605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.556605Z digest=sha256:eba0c4df8ae32dbc363d524f8e25718e8dd51360a74da098016cd9e59c8fee83

Observation 8038c5c9-bb5e-49fb-8d72-70ff9035f7cf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.560788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.560788Z digest=sha256:bd6ce0a5a963734224ac98b99ac3acef46ac5c62d131bd8a8edef70129cf48bf

Observation da7a92d6-6bd0-4504-8c55-5d2b3b2cf846 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.631020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:36.565094Z digest=sha256:e3a03d09f15b10ef1745cc4648e5b06a7e087ed96a88da6a1cc8a67e061f83fc

Observation f5069de9-453b-47c6-9d16-d3e7e0ac01d1 · outbound

This paper cites QuRating: Selecting High-Quality Data for Training Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection QuRating: Selecting High-Quality Data for Training Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.569987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.569987Z digest=sha256:ae037399ca74c0b0f0690c94046f4334c20211075566835fa338d527d62ead64

Observation 1648229f-8a32-450f-8ed0-2c1d03f64a6c · outbound

This paper cites Organize the Web: Constructing Domains Enhances Pre-Training Data Curation.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Organize the Web: Constructing Domains Enhances Pre-Training Data Curation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.576037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.576037Z digest=sha256:8426b0dbfe9451ca7d246db8148da0380160f9ffadbdaec16fda026564331664

Observation 7d1e7c72-9bd7-49a3-8403-c1dcfea521f5 · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.581532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.581532Z digest=sha256:b8112c5df0aa56dc4cf16f34e208d5a8be93278825318ecadd96ac5e1e85ffe6

Observation c1106b25-412c-4878-b6d7-a6f0c0abe522 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.Advances in Neural Information Processing Systems, 36:69798–69818, 2023.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Doremi: Optimizing data mixtures speeds up language model pretraining.Advances in Neural Information Processing Systems, 36:69798–69818, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.585807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.585807Z digest=sha256:88a30095ece1e793b0d15d18bc66680b63c39e19ec28877dc5ccae5422bf73fe

Observation 2105e29b-e248-4728-8d25-f58d390110a2 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.590686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.590686Z digest=sha256:23d5a0812207b64dcc023947ff0958efc2d04d8cfcc4ab71ed4f3228c218fe1c

Observation 9088e316-d5e9-46c9-b83e-f793a793c9f1 · outbound

This paper cites Mates: Model-aware data selection for efficient pretraining with data influence models.Advances in Neural Information Processing Systems, 37:108735–108759, 2024.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Mates: Model-aware data selection for efficient pretraining with data influence models.Advances in Neural Information Processing Systems, 37:108735–108759, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.608675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:36.596269Z digest=sha256:56249832321eb3a4ae8863cc3f6f702b281ea26064f00b080dd66a220a44d46c

Observation 727e857c-c636-488d-a5c8-3c70e7a52234 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.600908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.600908Z digest=sha256:9974a41a410ac8240a7a7bef3642e6bc27f1812d2d59f3f0647ebab656e4760c

Observation 67b5ae53-835e-4bfe-bd07-633422466e4b · outbound

This paper cites DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:23:36.894507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:36.606615Z digest=sha256:05a2056841c84afe8bf77b0c83b898b47d23b53a12cc82486d0ea2c13f03376c

Observation e409344f-83e1-4c8e-b326-ab841ba8882b · outbound

This paper cites Attention Heads of Large Language Models: A Survey.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Attention Heads of Large Language Models: A Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.611340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.611340Z digest=sha256:2bb4a4c7d76075f52c3b150cf19e4f950cd9cdae5f32725398f0284064f20e24

Observation 37fdae11-9ece-4c3e-95b9-a05642bae739 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.616305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.616305Z digest=sha256:43f8c99299ebcd5b30b58a12e6c723413c47b66e47d50287cbccb8be34836967

Observation 8c57fe57-fa0d-4e2a-afdb-205bbbc40cc4 · outbound

This paper cites Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts

Reference 55

Resolution
malformed identifier
no resolver link, observed 2026-08-15T22:23:36.648696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.648696Z digest=sha256:4c9b0a88c8ca49ac39ded46cc45fd0226330d015b342c87507a966916723dc63

Observation d61d900d-6071-44fc-825a-f19e204938d7 · outbound

This paper cites "This is a test string.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection "This is a test string

Reference 57

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T22:23:36.832878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:36.688883Z digest=sha256:42b5a623c5854ed754a0159f832308fa53e002db7d6cccb1fc28c520f53d4c16

Observation c016396b-fa74-44e1-822b-b2eef5aba8ad · outbound

This paper cites 26 Figure 18 The cloud maps of the data selected by AttentionInfluence and FineWeb-Edu Classifier, respectively.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection 26 Figure 18 The cloud maps of the data selected by AttentionInfluence and FineWeb-Edu Classifier, respectively

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.593637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:23:36.717737Z digest=sha256:f69530393de0a5e493e2f1cf13b1dc5a6a78235b136770b1a4c3f359e3b4d351

Observation 6fea8a0a-3d62-4e62-9dbb-15a270189090 · outbound

This paper cites an unresolved cited work.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.361541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.361541Z digest=sha256:96161bc6caaab5ab5bc089678d850ac25ca857615957408bd2f94b160a79f096

Pith citing papers

Observation 2ea950b2-81b0-4595-bed0-f55952963bc9 · inbound

Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality cites this paper.

Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.035631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T12:57:26.458109Z digest=sha256:054cc4f82da39c95fd82b8d522b6cf54868e4cc0eae34341a1e6a8291167cf42