Pith. sign in

Paper Citation Record · LEDGER

QuRating: Selecting High-Quality Data for Training Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2402.09739.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09739 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:10:55.615594Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:18:54.736613Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8f757b88-4eec-45e2-a1bf-77dd0261d419 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report QuRating: Selecting High-Quality Data for Training Language Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:44:38.290169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:20c22410f4f915ba513c1c1666b16738ff38b65263046824e6292f8e36966e98

Observation 8d2cf921-3ca6-4c58-97ce-1c58da76e6ec · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation QuRating: Selecting High-Quality Data for Training Language Models

Reference 287

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:06.815784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:915bfb9af1e054eb48e7670c253b0fcb0c37eb2aa552aa5ca229c6a7d15de708

Observation 412dddb6-e0a4-4dcc-9279-9fcf96d50a90 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models QuRating: Selecting High-Quality Data for Training Language Models

Reference 195

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.302685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:f5862b3d75f79ca1222951d1a8b9034cfe96920b6fede2ad923604d91429b76e

Observation a6d65f44-e795-43de-a1d8-c71c220a2dec · inbound

ProSec: Fortifying Code LLMs with Proactive Security Alignment cites this paper.

ProSec: Fortifying Code LLMs with Proactive Security Alignment QuRating: Selecting High-Quality Data for Training Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T17:10:55.615594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:10:55.615594Z digest=sha256:b9667235a09e375f4f70fc72ed60d3d1636754864db591e7caeeca4674f91bc7

Observation 1e9b78f1-6e8f-426f-88ca-371fd5b4f85f · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models QuRating: Selecting High-Quality Data for Training Language Models

Reference 211

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:02.085759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:02.085759Z digest=sha256:d76e19b0b266bebb8346c08c91684928eb491517eb11c59fb6b560db1dc8dfa6

Observation 1d999cc0-924f-401b-92b1-27c85c2d0ee2 · inbound

Weak-to-Strong Generalization Through the Data-Centric Lens cites this paper.

Weak-to-Strong Generalization Through the Data-Centric Lens QuRating: Selecting High-Quality Data for Training Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:28.530674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:06:28.530674Z digest=sha256:6b9df5be07971c2119e6bd5ccbaf10c39c7b092f6e46a3a1db33c2d1defc6623

Observation c3bccff9-ea00-48d3-a331-ff3b9315da16 · inbound

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights cites this paper.

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights QuRating: Selecting High-Quality Data for Training Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:02:00.883462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:02:00.883462Z digest=sha256:845503c7aaba5226d13b9e6e81dc2060b7125dbe905fe9483dc5979cd80f0881

Observation 44464480-61a9-4b3e-a21c-055ea8216edc · inbound

Optimizing Pretraining Data Mixtures with LLM-Estimated Utility cites this paper.

Optimizing Pretraining Data Mixtures with LLM-Estimated Utility QuRating: Selecting High-Quality Data for Training Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:05.091466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:05.091466Z digest=sha256:4198aa9df4c41c32188ae70bc92b9ad8c58bf3258fe0a0df909bde282dbc705e

Observation 55dfd1c6-7e30-49b7-ace6-f1c89c901824 · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts QuRating: Selecting High-Quality Data for Training Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.403659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.403659Z digest=sha256:4c3dcd2775ba4eecd4f28995a094161ada759549d0b164d7f745f3ce55e3449f

Observation f315cbc0-8d87-4f8f-81a5-ed74cc1a313f · inbound

Enhancing LLMs via High-Knowledge Data Selection cites this paper.

Enhancing LLMs via High-Knowledge Data Selection QuRating: Selecting High-Quality Data for Training Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.140161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:04.140161Z digest=sha256:d7ff9d9818b8752025c273ed537779cdbcf0d69f1a78a7d9a81b7b783ff9bf8d

Observation 1debfe81-31f3-4173-9bf5-783109667a66 · inbound

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis cites this paper.

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis QuRating: Selecting High-Quality Data for Training Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.179506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:35:47.179506Z digest=sha256:225df8015f8b2e8b126ff5cf68d559c52eb229fc7cc34b4b40907f81f51539da

Observation 290c994f-6f21-4700-9db7-cf29e9336f1d · inbound

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation cites this paper.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation QuRating: Selecting High-Quality Data for Training Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.115749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.115749Z digest=sha256:9ae4df18cdd417e4749b56642d4b39ae2767efdeab4c35347eb36d6a98709124

Observation 22547dd7-86c5-417b-91b3-1093554466e4 · inbound

LLM Data Selection and Utilization via Dynamic Bi-level Optimization cites this paper.

LLM Data Selection and Utilization via Dynamic Bi-level Optimization QuRating: Selecting High-Quality Data for Training Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:21:58.039374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:21:58.039374Z digest=sha256:a07667b9fe65cb6b43ca1bb0e2551f91e3c66f0fb49d4829f3f437f8134018e3

Observation f71f9222-a991-4346-b00f-6b81fd770bdf · inbound

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining cites this paper.

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining QuRating: Selecting High-Quality Data for Training Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T11:16:14.661866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:16:14.661866Z digest=sha256:8c95fb3f0fdfa14eba1c93bdb341c3b47680751d83d875e85600e264f5117081

Observation d20d0a03-a512-4361-babe-7be14fa90ff2 · inbound

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels cites this paper.

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels QuRating: Selecting High-Quality Data for Training Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:41:08.344219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T08:39:54.747656Z digest=sha256:7c86a32f3615c312e82539b4b134374b5a620a69480346a87ebc7cb9c77ed652

Observation 2ba2b4b1-181a-4c67-aad5-c3fe809936dd · inbound

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models cites this paper.

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models QuRating: Selecting High-Quality Data for Training Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:04.558118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:13:24.750244Z digest=sha256:80bedd659c86417be8edc16a81047dfc7044998254bbba024ae58989417e34bd

Observation 02bcdd64-91a9-4893-9093-8200f7e4eb2c · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code QuRating: Selecting High-Quality Data for Training Language Models

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:04.575591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:ca35f53bab07a9e3300dea97c3db7c262570b7c170d071cb00faffef3a4206bb

Observation 707067e8-c56a-445a-acae-50cf82b1f232 · inbound

Unified Data Selection for LLM Reasoning cites this paper.

Unified Data Selection for LLM Reasoning QuRating: Selecting High-Quality Data for Training Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:39.996586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-22T05:33:20.930156Z digest=sha256:7c911ccd35c64cea0f5bcd3202b6601a2978f4d9b333423fe78af2c29c89cf82

Observation 576fbebc-0fe0-4b8e-ab87-a90d1287d56c · inbound

DRIFT: Refining Instruction Data via On-Policy Data Attribution cites this paper.

DRIFT: Refining Instruction Data via On-Policy Data Attribution QuRating: Selecting High-Quality Data for Training Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:18:54.738254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:57:05.784589Z digest=sha256:0af07b56c6e4d86df67ecbde4df660f2950d03c734210319551b1aa5bef0f36f

Observation e8c41971-b1b2-4a4d-ab55-035f0385311d · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures QuRating: Selecting High-Quality Data for Training Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.419653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:9317011dd3cd40121a35db87bd60dc26fd885ce95ba38a569575ba814f49f3bd

Observation 8a4bba88-7a4f-4b5c-b9e0-987ea6a9ec2b · inbound

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators cites this paper.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators QuRating: Selecting High-Quality Data for Training Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.427294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.427294Z digest=sha256:2e608b7097da1aa955c013bb50c88a7d9cb6a6b5d7e91471ba68698e132a924d

Observation 4b8f4a6c-9809-4b7a-b3b9-2834d9d097df · inbound

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data cites this paper.

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data QuRating: Selecting High-Quality Data for Training Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T07:01:44.329991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:01:44.329991Z digest=sha256:25926f752f07ce398a4f0c8a2214728378839d8467640bdce8138d92d588a976