Pith. sign in

Paper Citation Record · LEDGER

MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.06046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06046 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:21:56.521904Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:48:39.405954Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f3e0ccfe-6668-4d09-b5f5-7a27ce0e6b6c · inbound

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs cites this paper.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.521904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.521904Z digest=sha256:3bc348ee4b73160f4873d25bd0fbd473e907aa77a051c6e4ed0e7fd3f82d6bd5

Observation c91ef90a-3734-44ab-832d-846d511e3051 · inbound

LLM Data Selection and Utilization via Dynamic Bi-level Optimization cites this paper.

LLM Data Selection and Utilization via Dynamic Bi-level Optimization MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:21:58.101569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:21:58.101569Z digest=sha256:53ba7e89c6db851322701c3f0cbf065b76953ec7ad394d905857af7c7e05ce5f

Observation 0ac98bfb-6693-4547-848f-59cec9b21172 · inbound

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining cites this paper.

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T11:16:15.281996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:16:15.281996Z digest=sha256:ad6557562bd450f3da2ebc30bb912d55ebd8ee98e7016e8aa8ee77baf9d5d336

Observation 668e404a-e66a-4f73-8486-a25ecaa60f98 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.804742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.804742Z digest=sha256:8adf1c2a55af3bc9374dba8c41d87e84eb8cf1cd4ef063770daccad62f0e0300

Observation 05bae7ed-8f67-43e1-87cf-22af4b88def0 · inbound

Efficient Dataset Selection for Continual Adaptation of Generative Recommenders cites this paper.

Efficient Dataset Selection for Continual Adaptation of Generative Recommenders MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:03.771646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:13:46.622289Z digest=sha256:08dd25ec3759b6c19d7d28bb781c7cabad0dfb9abe8d4daaaea6bf9ccacf4912

Observation d6b18352-4ab4-49ea-87fd-1b67c952ec14 · inbound

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models cites this paper.

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:04.521623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:13:24.750244Z digest=sha256:710047ae92944c10cc012bd29284b6b83cbb8a1e695f08781dd4b74e8ab3d332

Observation c1e8c891-7e22-489a-89ba-d09753bbf189 · inbound

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization cites this paper.

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 80

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T05:30:57.980667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:06:46.131725Z digest=sha256:d31a7e41f6fa4b995818e590fd9de4af79c3118794f0eb2f4ad6430b5ed898a8

Observation 030ce241-3d36-4241-873b-1a6b7690c8b4 · inbound

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation cites this paper.

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:48:00.969572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:47:36.122054Z digest=sha256:3dd86ff11510aa2435a0730d1b5f77b05bda6db52df2a530a5f59aeba1529e7d

Observation 1205b449-a336-4516-8c77-888553e55417 · inbound

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation cites this paper.

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T16:12:13.358123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:12:13.358123Z digest=sha256:b2904d2ac5da2333537c15a1ac5d57ee5e9cfbb9b840b9fdbd5e8e73d2b3d51e

Observation b3ced776-83bb-45e9-a0c6-bb99289f2fde · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.407393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:588d3e9a33acd48843b5335147ed16208fbff8f328b3e3fdab99e3239f696cd7