Pith. sign in

Paper Citation Record · LEDGER

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling

As of 12 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2501.13779.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13779 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:37:28.031565Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:30:13.859525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:30:14.503066Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57aa3840-cf97-490f-8bf0-3165d9225760 · outbound

This paper cites Data curation via joint example selection further accelerates multimodal learning.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Data curation via joint example selection further accelerates multimodal learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.987061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.987061Z digest=sha256:e822b1097cc6760bf13be180ddfa8caecc74af6489e9b30771b73731cb36d6f5

Observation 76e21ff7-a526-4546-8e6c-a79829391353 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.991160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.991160Z digest=sha256:5b9aa5547bf86776f8134f5f65564ad0dda9fe936616e57aaebfc0a6e35cb1d9

Observation 8f0e6cfd-f91a-4dca-ba43-7acdf99c0232 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Training Compute-Optimal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.995473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.995473Z digest=sha256:60c5038b5c762879afc51487572f11a2577a37d63744a3d80c0a22600b1a9761

Observation cf1cbce5-c727-467e-a0c7-3f0cbf45055c · outbound

This paper cites A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.007556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.007556Z digest=sha256:0ad3fadd413593606296d58d777caebaa5156753ceffe3c4aa77b8069bd165df

Observation df4be504-2cc3-45a3-8968-7a680def5720 · outbound

This paper cites Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.015364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.015364Z digest=sha256:05f6e18cb815f5504601516e7c4ab074ab1ea25803fa159289f58245ab36c339

Observation 493bc0ec-9dd7-4b54-aaff-a0627e07ee4e · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.019203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.019203Z digest=sha256:cc08c1c911254d8632db50ada4f63711c19674260dca46274e577f9a4923710d

Observation fa7f149c-1842-4da7-b1d9-d485e5ff0547 · outbound

This paper cites Topological Data Analysis Applications in Natural Language Processing: A Survey.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Topological Data Analysis Applications in Natural Language Processing: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.023924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.023924Z digest=sha256:f0a3942a9b9e56f2f948bff5c302c29fc246ebf36b6ddecd957423fbf454ddca

Observation 2a9d285f-a993-48d1-99f8-40abf99fe9a7 · outbound

This paper cites Distilling System 2 into System 1.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Distilling System 2 into System 1

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.031565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.031565Z digest=sha256:7b75c736a108c9f208c9d26215907f2e0dc8bc6817ff97c0a91b79eaae580b73

Observation 1d68e051-9436-4202-af41-b71856f84b7b · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.027916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.027916Z digest=sha256:c0a3ccb8160855ca7c9c902b8a791dd446dcac759cbc71a1a80ea00e8f41e5cd

Observation e72f662c-45d7-4076-836a-49b06127408c · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Quantifying Memorization Across Neural Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.978597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.978597Z digest=sha256:13d0e98481432ac1fe51aa5235c50c89823a6c33889d42e4ec50078cfff8d3f6

Observation 48e7baaa-e522-43e2-b2cc-0783039d3c92 · outbound

This paper cites Deduplicating Training Data Makes Language Models Better.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Deduplicating Training Data Makes Language Models Better

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.003718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.003718Z digest=sha256:f607999a91b33a0ac171370173c5548da12923f6545102b8c8b84c696f31f37e

Observation 27396786-c410-4a16-864d-89ac43ba7668 · outbound

This paper cites Scaling Laws for Neural Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Scaling Laws for Neural Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.999507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.999507Z digest=sha256:09fe633d8c45d6dec43dc8c3031cb21938358a748345626823ab9fe52ee25b4a

Observation de5fd0ff-ecc9-49c5-a025-124ef5306c59 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.011614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.011614Z digest=sha256:47c3ed19637820aad4dd2ad9366cedd9208c8c7332c92c2d6a47623f624990b9

Observation ab030787-fe67-4c77-9d86-88cae4a6308c · outbound

This paper cites The Llama 3 Herd of Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.983122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.983122Z digest=sha256:c425763845d1bbba77c849d5d2da4077f2631a8936f22197d87b7a6dc6b9a1db

Pith citing papers

Observation 4fc270a1-8595-4ca1-8907-7105aa6dcf02 · inbound

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training cites this paper.

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:30:14.508240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:30:13.859525Z digest=sha256:9bf155dcc996e134e8b64b70f530cb5bbc7b308d699374ef977054308f2e5442