Pith. sign in

Paper Citation Record · LEDGER

Investigating the Impact of Data Selection Strategies on Language Model Performance

As of 12 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2501.03826.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03826 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:49:32.290699Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 022ea559-4547-430a-9882-d87cf8325cd1 · outbound

This paper cites URL: " 'urlintro :=.

Investigating the Impact of Data Selection Strategies on Language Model Performance URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.219932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.219932Z digest=sha256:cfbf819f2b34ef7e10e5201238fdb53b1a304dd03e0404453683d6ed8755bfd8

Observation 5f32381a-64f2-49a7-81db-a394ff527064 · outbound

This paper cites write newline.

Investigating the Impact of Data Selection Strategies on Language Model Performance write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.225622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.225622Z digest=sha256:a63ef56fbfbc3aba2a437288423a300c20fd5ea99ad9c61cac1137dccf8fef09

Observation 21a5bb3e-e817-40a5-969e-8980509f53f2 · outbound

This paper cites A Survey on Data Selection for Language Models.

Investigating the Impact of Data Selection Strategies on Language Model Performance A Survey on Data Selection for Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.230472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.230472Z digest=sha256:5b0059b03d1b5090852256c074714e5691e4bced18c18d00503f72e38b33e2d4

Observation 3494af29-6327-40c2-bb2b-ed0d6a85f2fe · outbound

This paper cites Language Models are Few-Shot Learners.

Investigating the Impact of Data Selection Strategies on Language Model Performance Language Models are Few-Shot Learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.235627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.235627Z digest=sha256:2db327e7603a80f4067f63934bf2a78e0efc1d09aa0a3a6db4b44cda293279a6

Observation 239a1845-ca03-4a8a-957f-88f548584989 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Investigating the Impact of Data Selection Strategies on Language Model Performance Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.239877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.239877Z digest=sha256:d3fc085d7b0a3909f9abe390126d1ee3d9a68cb5c4ab941ab727cf37672b14ef

Observation 7a4fe3c5-57ca-4a87-bcc6-1536ca43ad24 · outbound

This paper cites Gradient-based Bi-level Optimization for Deep Learning: A Survey.

Investigating the Impact of Data Selection Strategies on Language Model Performance Gradient-based Bi-level Optimization for Deep Learning: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.245069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.245069Z digest=sha256:aaa32cc9e09c2764d61fc54e59b54942f846a045ab3bd6ed6529f87929ad5d4d

Observation 8d2d8e1a-85ae-44e4-b16a-1bde8b52f8f3 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.479451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.250611Z digest=sha256:d2e95ea8949ddfc1ae6e045c040982c4f8eff6525a54649bd1805323137afc35

Observation 9d9a4bc8-fc65-44a1-acd1-3cbbaeda5c5e · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Investigating the Impact of Data Selection Strategies on Language Model Performance PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.255398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.255398Z digest=sha256:84b0bbebdbc4f4ab7d8ba45ab993de5512d527994fc8c9e38c57f4d858abb4da

Observation eab36d1d-707f-402c-bf70-e98bedcfaa88 · outbound

This paper cites Don't Stop Pretraining: Adapt Language Models to Domains and Tasks.

Investigating the Impact of Data Selection Strategies on Language Model Performance Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.260531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.260531Z digest=sha256:a534be4513ee6ae2d68afc464a9a459edd863e04f8595623dcf484a43db8e9c3

Observation 3135496d-87e3-480f-a679-a225496398c1 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.266611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.266611Z digest=sha256:e0888e185ab72872125d158dae2b7780a164301e616d17b21b1ffca5e4bdec12

Observation 8c41f9fc-4dfc-4113-a92f-3022d9c015a0 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Investigating the Impact of Data Selection Strategies on Language Model Performance Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.271065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.271065Z digest=sha256:28a7aee4a1defc9a6bd0573a69aea8af168751f339215cbbdcfd699d00040460

Observation 042a6e48-3a63-452b-be60-f6e89ccf26f6 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.459945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.275387Z digest=sha256:bac72b22fc7f39ca982fac58ac6eb6dad96f4da0ab54cf4e19b4f0e4d8ce6c58

Observation 326e4e8e-f9d4-4b96-948f-5081f9855e58 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Investigating the Impact of Data Selection Strategies on Language Model Performance GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.278940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.278940Z digest=sha256:11d48b79c1cf07ad3a8cc966019c42d966f07a1a127e7617b2e8ba2820f8f4e2

Observation f520713b-44c9-4956-994a-6a7ff46e5da7 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.446891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.282966Z digest=sha256:e6e9aad999ecd983b9192c7326a7de9b8110b7cbc395b6ff702f6630df638105

Observation 0a93cd73-86ba-4d1e-9729-1e33082a474d · outbound

This paper cites Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method.

Investigating the Impact of Data Selection Strategies on Language Model Performance Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.286703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.286703Z digest=sha256:4064d9eba32b424ebc75a430530106ca8c72669a30d3b28f169bfadb834be6a2

Observation 3ee829c9-ff87-488b-9849-4fb2ba4547c7 · outbound

This paper cites Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities.

Investigating the Impact of Data Selection Strategies on Language Model Performance Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.290699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.290699Z digest=sha256:58d91e375a5f30b129db60c91a1763de8c592cc571665614564208c00d5a7a4a

Pith citing papers

No inbound Pith citation observations are available.