Pith. sign in

Paper Citation Record · LEDGER

How to Train Data-Efficient LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2402.09668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09668 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:58:34.553634Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3d59eb21-f235-4807-978f-cb47c3275d11 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models How to Train Data-Efficient LLMs

Reference 237

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.315894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:6f1b1a7b99658b33d65da13e50e8d5bd5308c3313613dc0b1ab8c486d8c22d80

Observation 792957dc-e756-4ca0-bad0-21339c6cd949 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:44:38.337372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:1db3a22b8b627f898504347c3f451f8d17a1511ab92aa48744173ef32f5d7333

Observation 149a544f-2fdb-4df2-bf3a-ca4639ca9b4b · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models How to Train Data-Efficient LLMs

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.197578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:27fc280463c1fa824c530530de40833f746115489997c5c21d66a3d5f3441e9a

Observation fd3cbe97-8165-41a3-9aa7-4928f63c3f7a · inbound

FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy cites this paper.

FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy How to Train Data-Efficient LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T18:58:34.553634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:58:34.553634Z digest=sha256:773376e62ff87b79526b29184fb6099cbf1f827012e26ae5d6b0f76df7a34936

Observation 1d43c75f-11a3-48e7-8928-a35bf4795c4b · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts How to Train Data-Efficient LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.347017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.347017Z digest=sha256:6d9e45225b9620825db3f47e23fb8c1ab55f8d58f77e1fb34b3ba129ca6c2dcf

Observation 2e7b6d1c-21d2-49e5-8281-2fa89251ff4d · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report How to Train Data-Efficient LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:22:12.189477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:5d9ae6e7bf487768593b871ddd962a2f28527fd2a5ff3c651317ee0f4bc0a458

Observation 008d596f-f874-4fec-ae02-f2b8d97f73f7 · inbound

Enhancing LLMs via High-Knowledge Data Selection cites this paper.

Enhancing LLMs via High-Knowledge Data Selection How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.600455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:03.600455Z digest=sha256:001feedc98ae0db1fad0b4155ddfd9c0e150ebffaae1231b61c4ee90be02141f

Observation 75691a37-0af0-4eb6-8cbb-6578d75e4297 · inbound

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain cites this paper.

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:54.356737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:54.356737Z digest=sha256:2270abb2f0ef37568b64ecf3d5991086e111f0e34b6bce5a4f1d724538649556

Observation 7be11a93-fc4f-456a-b1ce-c09431db1d5d · inbound

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining cites this paper.

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining How to Train Data-Efficient LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:49.740308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:09:49.740308Z digest=sha256:ff4d9eba259b7fd029689767fdc63123dd83fb362157b98aee0a1c05b44c3560

Observation 8e5f661e-1df7-4c21-b0f1-6e0f8d337875 · inbound

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets cites this paper.

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets How to Train Data-Efficient LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:21.726501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:21.726501Z digest=sha256:96afbefff5d06bbf4285834c9e552db6ffee94700b264387aed6b2a460bf9759

Observation 45a2bb00-27e7-413c-aad1-04bfb0365e04 · inbound

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models cites this paper.

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models How to Train Data-Efficient LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.769180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:19:31.769180Z digest=sha256:b196e0c4b72bc13a9cbd049ad0e23a057edf78abf0b9d9b4c1af81b7ef857e56

Observation baf9518d-8283-4d8e-9693-627a73a6bf4c · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning How to Train Data-Efficient LLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:22.198008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:22.198008Z digest=sha256:95e3f3d76a438ede2685d5f538fb0782d983808b6e3fcdfdc01ab61f4b888e41

Observation 9d290ed7-dc77-45a3-b4e6-29d20b1418fa · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation How to Train Data-Efficient LLMs

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.750924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.750924Z digest=sha256:5ccd422cbbdef580f99d941b2ee7424e30a07e4d855b3a559ed7960d776b9f43

Observation 2a1f4957-80ec-4cd3-a2fd-c1a72c8777ca · inbound

Assessing the Role of Data Quality in Training Bilingual Language Models cites this paper.

Assessing the Role of Data Quality in Training Bilingual Language Models How to Train Data-Efficient LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:57.821522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:57.821522Z digest=sha256:a85f9710f32a83a197a2631fed0dbfdb750d4c30928ea1a10dbfcca7df6e2c4f

Observation 98bac85f-4f08-4101-acf8-8a2b1e0ef229 · inbound

Disentangling the Roles of Representation and Selection in Data Pruning cites this paper.

Disentangling the Roles of Representation and Selection in Data Pruning How to Train Data-Efficient LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:24.716339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:15:24.716339Z digest=sha256:6533db6d13d7b214bc89a7c4d45f19aaa4b202ca06656f75c4f472cefc07a080

Observation f1040ee3-a77b-4e43-8c1c-ea8522bbea7f · inbound

Efficient Training of Deep Networks using Guided Spectral Data Selection: A Step Toward Learning What You Need cites this paper.

Efficient Training of Deep Networks using Guided Spectral Data Selection: A Step Toward Learning What You Need How to Train Data-Efficient LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:46.325370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:46.325370Z digest=sha256:2ef81f57b912a053c5b7b35fabb5433d050354ed1af7a10757e6f3f3a1d24d5a

Observation 6d867eea-4249-4856-8592-b3db67339c98 · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs How to Train Data-Efficient LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:44.081602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:56:44.081602Z digest=sha256:72b5b1e81a9e33de709d34dcb9fe2d22f377c9d60c3b4b0b34af5e6a5d8c0bdb

Observation d6c6eb5c-08c4-4a4b-a6c8-30527e315b02 · inbound

Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification cites this paper.

Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification How to Train Data-Efficient LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:32.871288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:21:32.871288Z digest=sha256:97e8d175cfbedef02ce3418f6ac03a07394810695a2620917a8c10b0b4ee0478

Observation 8490cd5d-1c4f-4bab-95f1-f8eca72731b9 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks How to Train Data-Efficient LLMs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:13.006591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:13.006591Z digest=sha256:81cead57c816cb873cf6d5030943b65cb19d5a1bf7c6c8f24d3cb6d75f355cdc

Observation 3a86b5a3-61a3-4750-a132-2cb73815b823 · inbound

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection cites this paper.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection How to Train Data-Efficient LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:42.206386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:42.206386Z digest=sha256:358ac6a35d54a90af44761fab20353f561bf19971f3b971378663758dc1a25e6

Observation 7860f5b1-4c1e-4780-86e1-89b7fa135eb3 · inbound

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models cites this paper.

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models How to Train Data-Efficient LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:04.764330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:13:24.750244Z digest=sha256:f3d9e2387be2c98f79856ba27ef4965a156be13ffe16ed257bfb9f9c9a12a74f

Observation 6b912b53-402d-4d00-bb74-0e15c36c04ac · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:59.080239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:0826c4b203c15346f01b78d0ee2024cd57737dce10c799bdd99c90a4223b9811

Observation 2b4a269f-1cbb-4eaf-b0d9-3a16f4405fe2 · inbound

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates cites this paper.

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates How to Train Data-Efficient LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:41:04.697353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:22:59.883307Z digest=sha256:c3a5deed1f66ce431fe55973cdec7add9e72629af3e94b193138786139c68a8d

Observation 6601df6f-9c38-41f9-ab46-d3096b43ba4e · inbound

DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models cites this paper.

DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models How to Train Data-Efficient LLMs

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:21:26.951063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:19:11.231416Z digest=sha256:1e0d08b3e07d5e536fc6ce4a5bba29b36cc8f2943e34291afc206a5378ec919b

Observation 98535e94-9a8d-4b8a-8662-34780db3a8c3 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:08.529589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:6e2a33f400c9a294048ccc3885de52d3bd663a62ee9b3bb3f09987f8dd87e7f1

Observation af26c0cc-ad3e-43fe-9683-2353c9ae1b0d · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:54.594603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:623b1bfed21d45eb482ff98b48a77a90b46dff8de058c13746f447e23cfa4a77

Observation 9fe4cc44-6781-46ce-ab2f-60a76e197dfd · inbound

Accelerated Relax-and-Round for Concave Coverage Problems cites this paper.

Accelerated Relax-and-Round for Concave Coverage Problems How to Train Data-Efficient LLMs

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:05:57.094895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:52:41.972226Z digest=sha256:51067791daf631a50e472dee40e8fa5ea3f677e2f1394b59e1668e5b44ee6e8c

Observation ad0f918b-ca1e-4b43-a0a9-df8466e9b472 · inbound

Reflections and New Directions for Human-Centered Large Language Models cites this paper.

Reflections and New Directions for Human-Centered Large Language Models How to Train Data-Efficient LLMs

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:25:56.843634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:25:46.378350Z digest=sha256:3539d92b2a2b69e2f5e342cc0b40e4f53693ada1624c34415f5d8addba7281bd

Observation 10576518-8575-43d8-9ea9-7263c128c47b · inbound

Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching cites this paper.

Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching How to Train Data-Efficient LLMs

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:03:15.950487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:58:52.511363Z digest=sha256:6acca0f9edc1b7c7aa53f6754bbfa5f673410be2c5a72e2418133ccc199b4535

Observation 741108f7-c40c-483a-9b45-dd06febbe00e · inbound

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them cites this paper.

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them How to Train Data-Efficient LLMs

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:22:46.237668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:19:22.355753Z digest=sha256:0a118cc8dfecd00fcc98266e7a49107829ae17680da3db82a7da918ba6cdaf94