Pith. sign in

Paper Citation Record · LEDGER

Investigating the Impact of Data Selection Strategies on Language Model Performance

As of 11 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2501.03826.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03826 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:49:32.290699Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 022ea559-4547-430a-9882-d87cf8325cd1 · outbound

This paper cites URL: " 'urlintro :=.

Investigating the Impact of Data Selection Strategies on Language Model Performance URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.219932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.219932Z digest=sha256:fb44d142ccfa0670f9bec0f689e52dcccf24ac89d91d71b3852463104d267a00

Observation 5f32381a-64f2-49a7-81db-a394ff527064 · outbound

This paper cites write newline.

Investigating the Impact of Data Selection Strategies on Language Model Performance write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.225622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.225622Z digest=sha256:8e2483b3caf06b78bb37c6ee6bdb4aa529c8eed8db84f023c1a4e767f38e10e3

Observation 21a5bb3e-e817-40a5-969e-8980509f53f2 · outbound

This paper cites A Survey on Data Selection for Language Models.

Investigating the Impact of Data Selection Strategies on Language Model Performance A Survey on Data Selection for Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.230472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.230472Z digest=sha256:c5c3afb0ed64a2e3450b942f5d7b37637f3e837900d8a204de4061c8280c5fb0

Observation 3494af29-6327-40c2-bb2b-ed0d6a85f2fe · outbound

This paper cites Language Models are Few-Shot Learners.

Investigating the Impact of Data Selection Strategies on Language Model Performance Language Models are Few-Shot Learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.235627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.235627Z digest=sha256:7fe17addb6d90bc8f1ae94acb6ddb6fc1f3aacb336fc782894b1d00d78217226

Observation 239a1845-ca03-4a8a-957f-88f548584989 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Investigating the Impact of Data Selection Strategies on Language Model Performance Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.239877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.239877Z digest=sha256:a2d491a2efa9d8d3f538f5441f967321dc72da89ed7876326019546c50ac40f7

Observation 7a4fe3c5-57ca-4a87-bcc6-1536ca43ad24 · outbound

This paper cites Gradient-based Bi-level Optimization for Deep Learning: A Survey.

Investigating the Impact of Data Selection Strategies on Language Model Performance Gradient-based Bi-level Optimization for Deep Learning: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.245069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.245069Z digest=sha256:b6b604200701504c7d50555430ad01c61708d9d087b31eb03941be4ead91d3d6

Observation 8d2d8e1a-85ae-44e4-b16a-1bde8b52f8f3 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.479451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.250611Z digest=sha256:0b1f10e1e69454d8c1f5b9e606f8944e25374f4cb15d623b486124337739d1d2

Observation 9d9a4bc8-fc65-44a1-acd1-3cbbaeda5c5e · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Investigating the Impact of Data Selection Strategies on Language Model Performance PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.255398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.255398Z digest=sha256:b9601853953e093e11a11d85e866baf9235aab6a0b7c630791b038cfd87659c8

Observation eab36d1d-707f-402c-bf70-e98bedcfaa88 · outbound

This paper cites Don't Stop Pretraining: Adapt Language Models to Domains and Tasks.

Investigating the Impact of Data Selection Strategies on Language Model Performance Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.260531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.260531Z digest=sha256:31076f7a71d506fbd7c9f37a09037bcaf851d079e5735515c0aebc4e08a53116

Observation 3135496d-87e3-480f-a679-a225496398c1 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.266611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.266611Z digest=sha256:9ab348912633dfa666103afb145fd873c002e17bb719b873d53f09fb8bd02bdf

Observation 8c41f9fc-4dfc-4113-a92f-3022d9c015a0 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Investigating the Impact of Data Selection Strategies on Language Model Performance Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.271065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.271065Z digest=sha256:723447a5fb4a6d248ad5083edbe5a280286787c63ce5f53a85caa35659301960

Observation 042a6e48-3a63-452b-be60-f6e89ccf26f6 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.459945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.275387Z digest=sha256:128554932f9ee4b7be6888b983e344ecb3f550b7bb9c604baaac28c68ebd89bb

Observation 326e4e8e-f9d4-4b96-948f-5081f9855e58 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Investigating the Impact of Data Selection Strategies on Language Model Performance GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.278940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.278940Z digest=sha256:a82b28f6dd51db1205a6642e62c9734f7ba33d5539d8bf14269e519c4a9903a6

Observation f520713b-44c9-4956-994a-6a7ff46e5da7 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.446891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.282966Z digest=sha256:71e42f59fe44d672195074ed9824050a3c3cca80d5de0e1ed67444f769ec37eb

Observation 0a93cd73-86ba-4d1e-9729-1e33082a474d · outbound

This paper cites Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method.

Investigating the Impact of Data Selection Strategies on Language Model Performance Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.286703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.286703Z digest=sha256:b851d3210003d396eb517792bb211fd3ac99ca2fdcd763ea564f6d0816c7f6aa

Observation 3ee829c9-ff87-488b-9849-4fb2ba4547c7 · outbound

This paper cites Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities.

Investigating the Impact of Data Selection Strategies on Language Model Performance Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.290699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.290699Z digest=sha256:7fcdc80f5b70eee568e8b51ebc44b0d5d8ba02ad8e1d9030b7add36a1158e4fb

Pith citing papers

No inbound Pith citation observations are available.