Pith. sign in

Paper Citation Record · LEDGER

Predictive Data Selection: The Data That Predicts Is the Data That Teaches

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2503.00808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.00808 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:58.609898Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:47:05.647227Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da652bd2-770b-44a9-9884-0a3d163566cd · inbound

Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training cites this paper.

Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:58.609898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:58.609898Z digest=sha256:c233504e93c40f9186ba48c78cad49ab3dcf0598db71f731396093dbd63d86aa

Observation 17216e23-cf4a-4cb2-8bc2-680b7e14c519 · inbound

BlueLM-2.5-3B Technical Report cites this paper.

BlueLM-2.5-3B Technical Report Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:55.930316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:20:55.930316Z digest=sha256:9918f65467aba400c276ba64825a4a2eae63316b3ff03e181fa8ecbc8bfea224

Observation 7183ee2b-8b22-4e68-bad4-1f69c26df013 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:13.692690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:13.692690Z digest=sha256:081da3f25c53a9fe841cc59ecf9505634939d23fb8965f1f3540ef87132d780d

Observation 04727587-5435-47c6-9a54-2a8a9b07f2fd · inbound

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation cites this paper.

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:06.812208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:06.812208Z digest=sha256:a1f3f515d5d7cc923503e29e4ce4f8ef8fc280be7f31da3b5bdfd7febb277e37

Observation ce4c0162-ad61-47b9-8de7-a77b68c24dc5 · inbound

Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis cites this paper.

Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T16:57:07.095586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:57:07.095586Z digest=sha256:2e968b50ad903eab9ca5fb0df1f20ce17a27f7b3bd532f6c357d691475cd60f2

Observation 2237c73b-6452-4317-a938-f70c0f8f7c9c · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:18.455452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:24d5b1d0d445e0af2649e4646fb4ce92ff7fe786925465b4a41f203a4b6928a3

Observation a8eaf7d2-9d56-4e53-99bc-0e1d8eff1275 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:54.537057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:3f62dd5969235ac56c525a758758587832005d54d1781a94e245cc9a54ecd1e1

Observation ee7886de-bc9a-43f1-a59a-17a12ff0518d · inbound

CausalMix: Data Mixture as Causal Inference for Language Model Training cites this paper.

CausalMix: Data Mixture as Causal Inference for Language Model Training Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:47:05.648948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-02T15:45:02.415577Z digest=sha256:3bbafdd99985418026486a97d7d3bf9f1ac47e1a509a0c2e8c94bb2543863a20