Pith. sign in

Paper Citation Record · LEDGER

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

As of 23 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.04969.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04969 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T10:45:46.618668Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a1d000a-3215-432f-be92-848e41b96587 · outbound

This paper cites Phi-4 Technical Report.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:0b658b3f26f87893b629e170b77562f89072409135e2c5d0430fc90edaaeb756

Observation 21cff844-1be4-4e10-bcde-fd3fc1f8428f · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:6a78a813c0c569bb824b4f32bb5aa5d3bd1927eba438eaf0baff68f324f1df3c

Observation 57b7f3bd-d63e-46e9-8f64-c6b2c8044616 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:1e0104c6bac891b757a5910bb1f15d7b7667685eefee13124f0cf93256961996

Observation ac4162bd-e12c-4948-a030-39ba601e5c8c · outbound

This paper cites Reformulation for Pretraining Data Augmentation.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Reformulation for Pretraining Data Augmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:83476b2cd3cebabe22eb31a7a210cc8005d07115a01346ca8d9755b616919f94

Observation 0f37e195-1eca-4716-a8b1-1c6d23669b05 · outbound

This paper cites Scaling Laws and Interpretability of Learning from Repeated Data.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Scaling Laws and Interpretability of Learning from Repeated Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:8b947976064830468b570f98034097d4b5c85da4571ed63b495abbbf36af10fa

Observation 0736a376-dae2-4317-820f-6f4f87747bc7 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Training Compute-Optimal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:4fb6b1079c3c2eb08b369dfcf05c81fc68a37bc322a10ff121107f8af93df665

Observation e364a06c-7f2d-4cab-aeaf-bc5df9d406ab · outbound

This paper cites Ultra-Sparse Memory Network.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Ultra-Sparse Memory Network

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:5e1b62034f7608c87024a0263e11e252ff5eeb020eba429eaa99d378b22b9e0d

Observation 5b965e8a-d04d-415a-a662-6fe5946470a2 · outbound

This paper cites Ziyue Li, Chenrui Fan, and Tianyi Zhou.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Ziyue Li, Chenrui Fan, and Tianyi Zhou

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:36fe57544c1fbf7019bc7a05082277a80aaab6c3136fa6de31789ed9399025ad

Observation d0dfc0bd-c81c-49e7-a214-27aa69869332 · outbound

This paper cites Let's Verify Step by Step.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Let's Verify Step by Step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:032959105d5ad7944cfa7c52b5a976908bb0362fe424df097f59f22b2a08d66f

Observation 7ddfe9df-2dac-499f-83e3-0f48e186728f · outbound

This paper cites How much do language models memorize?.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training How much do language models memorize?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:7e8dc8dfcc0fcdd324a303707dbdaa981b65dbd144165e3bd27e98dbe643cc4e

Observation 016b0d27-4f1c-4fdf-9077-d0a313695566 · outbound

This paper cites Olmo 3.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Olmo 3

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:14f2adb0647b4565fee6b2b10f01edeb1ec6eae7b6027ff4b84dc2294b9198c3

Observation 5fb69b54-f28c-4bea-95ec-b9069f16e924 · outbound

This paper cites Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:ab7c919ac238226bd40c47ba5db90afdf1c86c8d34ced3f655201428e2a9b663

Observation 9d5df397-5493-4268-bdaa-aa093e709d94 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:655ca9049c3efa9f48c7a2314a358a3305e264a3af000ac8231754409c7c31fc

Observation 9f79dfb9-1d36-46f9-a59d-b50b33d7fb36 · outbound

This paper cites Galactica: A Large Language Model for Science.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Galactica: A Large Language Model for Science

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:825ac91225e4887d4b7ed7746e920fb4c1989c69ba19f78a69dbf3409478588f

Observation c2f585ae-f956-40fb-ba97-e8247b0a8355 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Kimi K2: Open Agentic Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:814bed6100b9c1161bd7c42cd0b0d29d75696fe010b3f86919ce70f651666d55

Observation 72e0f637-ddf4-4168-bb56-6415d1dfcb4d · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:ef0fdf66764b42e4ad48cb7b492330c9b199bf80f03e3d1c0cbc1d663783d9c1

Observation 10a4e715-c655-4ca4-aee8-55fcf3ac4764 · outbound

This paper cites Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:6d0ce29495d97acdc430786b973db0057cc2be3286c94838d1780e42fe0f9109

Observation d13ccbe8-d9a3-40c6-90d6-4c078dd88cd3 · outbound

This paper cites Larger datasets can be repeated more: A theoretical analysis of multi-epoch scaling in linear regression.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Larger datasets can be repeated more: A theoretical analysis of multi-epoch scaling in linear regression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:74af26fcc128799018747670b530b23a8d1e685309a4f94c2795838aba4bcb94

Observation 920c41c3-595b-4484-8567-053d0cb39d89 · outbound

This paper cites Nicolas Zucchet, Francesco d’Angelo, Andrew K Lampinen, and Stephanie CY Chan.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Nicolas Zucchet, Francesco d’Angelo, Andrew K Lampinen, and Stephanie CY Chan

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:995dd9d28a80f0dd97de322088fc96a0db40f49de98b4c721644eefaf95f80d2

Observation dff1143d-d849-420a-9631-c9316323ac43 · outbound

This paper cites Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance.

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T10:45:46.618668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:45:46.618668Z digest=sha256:25434674ad3f6e71bff16c5e7b4619a96b6214e82821a1fddb389a37fd2f889d

Pith citing papers

No inbound Pith citation observations are available.