Pith. sign in

Paper Citation Record · LEDGER

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

As of 14 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2506.19262.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19262 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:19.055117Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:53:39.113934Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy6
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 645864b0-6078-460d-894c-0f73a24b1c66 · outbound

This paper cites Phi-4 Technical Report.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.851748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.851748Z digest=sha256:5cc7e2aad7853ed94c4bacb34a50132c6298d3f93b37cb736cfd76a56c538a7a

Observation bba064c8-de37-4c7e-84c8-f3930b16b45b · outbound

This paper cites GPT-4 Technical Report.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.857973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.857973Z digest=sha256:78108872e613fbe98c90e29754e8fbdc472aaa1bf0543a8aac8e3db4f72a4117

Observation 7043876b-332c-475d-a4fb-22d500c877cd · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.863690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.863690Z digest=sha256:24615bbcc5a44bc98607e0e096a5c38e38b84c35935e47b32630fa393a3b2287

Observation 6492b30d-3ec5-4f39-a21b-099f9cb35650 · outbound

This paper cites Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.868972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.868972Z digest=sha256:f19be3ca3ccb889a7480ecb3afc5c6c93c6cc364c901dc38c1730efac56fdf80

Observation b6c58840-c707-4615-bd4a-26565c4b61e0 · outbound

This paper cites Language GANs Falling Short.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Language GANs Falling Short

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.875688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.875688Z digest=sha256:7f53f55ffc38a7bc4a2741152bbe9216d540226cdd33b761be5f5e4e9d13d3bf

Observation 89491718-442f-463f-aaa7-146c097564c0 · outbound

This paper cites On the Diversity of Synthetic Data and its Impact on Training Large Language Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning On the Diversity of Synthetic Data and its Impact on Training Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.881689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.881689Z digest=sha256:d2a0ef15ee3d663e24cf242ca7a647eb5565c714d94ee5fd859f26c8c00635ae

Observation 5567f7ba-0961-48e8-9080-1a02a7cf09ed · outbound

This paper cites Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.888836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.888836Z digest=sha256:15bac10cb5977c591145bce8bdc9f9be281157680a105ba6fc3e3b09d6d859d7

Observation 505260d3-412a-48a7-a24d-6128c7fccdf9 · outbound

This paper cites AugGPT: Leveraging ChatGPT for Text Data Augmentation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning AugGPT: Leveraging ChatGPT for Text Data Augmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.893945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.893945Z digest=sha256:540dd0350c24b569f33f9cc19e6e3df68885ed333980d3ddbf9d2d2890bb0808

Observation 0df68e44-eaa5-4ce3-a3f6-002123cd2fe3 · outbound

This paper cites Universality of the π2/6 pathway in avoiding model collapse, 2024.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Universality of the π2/6 pathway in avoiding model collapse, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.787288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:18.899052Z digest=sha256:2653e9b3c10211e884e22a6b6236812e89e08c85087c60031ba8cf3b526505bb

Observation 42019cdb-260c-4c0f-a28c-833ec6c747c1 · outbound

This paper cites Is GPT-3 a Good Data Annotator?.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Is GPT-3 a Good Data Annotator?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.904149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.904149Z digest=sha256:27e8054da9b65fe2a0d49d8eb168d6218184034fabeb5657cc0067a6f3b36803

Observation 8c071fe3-c614-4553-ba3d-e2f9dc40d8b9 · outbound

This paper cites Data augmentation using llms: Data perspectives, learning paradigms and challenges.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Data augmentation using llms: Data perspectives, learning paradigms and challenges

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.769259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:18.909378Z digest=sha256:c30b036d80243f382b6f1380bc6e0d9af7ed53bc4168745dac23bd84ce4c3b4c

Observation e53f2ab7-2b3f-4ea7-98ae-47bb7a5f405d · outbound

This paper cites Model Collapse Demystified: The Case of Regression.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Model Collapse Demystified: The Case of Regression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.914754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.914754Z digest=sha256:fc009b79eb4376d69e433d62f38e262d42a7c96673a5da69cc6a90441907b9d1

Observation 5bcb2006-a0a3-4586-9a10-828463c2288e · outbound

This paper cites Strong Model Collapse.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Strong Model Collapse

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.920822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.920822Z digest=sha256:97b773f4ecc5694272ae826da8cf63a462ae64c08e738b0993eef1d566e4a69a

Observation fd1d13bd-5531-4d84-ac15-50c6891f3cbe · outbound

This paper cites A Tale of Tails: Model Collapse as a Change of Scaling Laws.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning A Tale of Tails: Model Collapse as a Change of Scaling Laws

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.925506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.925506Z digest=sha256:fe8dcf00334a40dd31841b651e06d822affb5c42413c5f97b33a670c0bab70d7

Observation c0e572d1-5b86-4daa-95cb-0c81f392a9c2 · outbound

This paper cites The Llama 3 Herd of Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.930500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.930500Z digest=sha256:d6a2ddc660329955b9bbb94d33f70cce6683030f848b99f7491b09273bb82796

Observation 4a98b15e-7d2e-4c28-b846-013385ff65ff · outbound

This paper cites Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.934682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.934682Z digest=sha256:ccb6d25fa81eaf499f95e7dce06ef967692036be285805bcb873ed56a2e00dcd

Observation a81f5911-b703-410c-8c4a-ea02c64c4425 · outbound

This paper cites Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.939264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.939264Z digest=sha256:7468cacdd41969f0681ea4fb89a1a5c0915cfcdf80fb10ed89465b674d58b5c3

Observation 1991e3be-c794-4a86-925d-05e592661c07 · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.943859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.943859Z digest=sha256:92055ac064aae7b4242d0a114f7a0ef0844c76dbe1974d1cd0f9e6658cb281e4

Observation d43b8f4d-bb3b-462c-8989-9ac16ce16723 · outbound

This paper cites The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.948737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.948737Z digest=sha256:a4be13297ab3901f309168c803a593adac863268a5e9b2acc1f9c086d5430374

Observation e7322882-a784-449e-93f6-5e77106b91ab · outbound

This paper cites TarGEN: Targeted Data Generation with Large Language Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning TarGEN: Targeted Data Generation with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.953175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.953175Z digest=sha256:90b8b3f99840fb40983d91104ec8d80a8dc36fe90e9ba4b3c5080d68c7c2c0fc

Observation 7de1fa9a-a33c-4f24-8c98-16f7edaedbb3 · outbound

This paper cites Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.957552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.957552Z digest=sha256:9c45b874b1f8aa95fe56af6170a48f725b60cfc91bcca18d6f8c64af1e9dc5eb

Observation 551809de-a8ab-45bf-a7ed-46802623ffbc · outbound

This paper cites The narrativeqa reading comprehension challenge, 2017.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The narrativeqa reading comprehension challenge, 2017

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.962396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.962396Z digest=sha256:c915d51e55fb85cd61f9f9d12dae3c6210c727356e261afeef3c9858093c4e2b

Observation 00dbaa3d-0fe8-48ec-ad14-e4843330d71f · outbound

This paper cites Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:19.334959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:18.967285Z digest=sha256:92efdff2cc6d200f0f6c75e5b60c7645395e0b8f4d4e9599a077725383af28ad

Observation 7c60ce21-922e-463b-afa8-1cbbd6de9982 · outbound

This paper cites A diversity- promoting objective function for neural conversation models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning A diversity- promoting objective function for neural conversation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.732046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:18.972529Z digest=sha256:888eebd354fbcc2767720ac36aeb7f20c56d0189d909a6cdacb9c23e4cea4447

Observation 5c87fad9-1675-4c54-a5c8-50224cb584aa · outbound

This paper cites Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.977169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.977169Z digest=sha256:5f26dff85577adaf314b83cee2fff47b62fa40f3bd736b25253d973d84161028

Observation a74c14c8-a7b2-4817-a577-3595866ff611 · outbound

This paper cites API-guided Dataset Synthesis to Finetune Large Code Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning API-guided Dataset Synthesis to Finetune Large Code Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:19.293347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:18.981774Z digest=sha256:b7b3a0ebe110312ee9ae1d676bc71882331e371972dd5ecdf0c1f06341e5b6bf

Observation ec65d921-5081-405f-910d-cd222a216ae7 · outbound

This paper cites Generating training data with language models: Towards zero-shot language understanding.Advances in Neural Information Processing Systems, 35:462–477, 2022.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Generating training data with language models: Towards zero-shot language understanding.Advances in Neural Information Processing Systems, 35:462–477, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.716531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:18.986716Z digest=sha256:d68e31c273ed50a4d8a8c397d0e96d499d2335bff447ed3bb9c944575e420407

Observation 04faf82d-35c8-4fb2-95f5-f385cf7333e1 · outbound

This paper cites A corpus and cloze evaluation for deeper understanding of commonsense stories.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning A corpus and cloze evaluation for deeper understanding of commonsense stories

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.990942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.990942Z digest=sha256:7e5bad00474680ca66699df2bba9253297529c0830bd8d98fa7c7cd70562ef4a

Observation e8b62230-41bd-4f64-ad14-816ffc6addee · outbound

This paper cites I learn better if you speak my language: Understanding the superior performance of fine-tuning large language models with llm-generated responses.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning I learn better if you speak my language: Understanding the superior performance of fine-tuning large language models with llm-generated responses

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.685719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:18.995370Z digest=sha256:4b5fa1b8a530e981fca11eaf7c16bc872e08e2cd71577822122a931f17de714b

Observation 9e77a416-a27a-4afa-88bf-749bc2797c13 · outbound

This paper cites How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.999831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.999831Z digest=sha256:03e8ddee65ea7dfe62d97af1a3cb536e47842ddb0be121b862370d1342d73b55

Observation 7562b1ae-da84-49a0-98c6-5009f362cd28 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.004646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.004646Z digest=sha256:bc4140e64548a9cba5750ad81e63a8e9e77d64b9c77e7ba43ff3b8d7826bc89c

Observation b8ed6bdd-5545-402b-8ab4-7adc93da7581 · outbound

This paper cites Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.009713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.009713Z digest=sha256:899e02849d2ab26d1d9b59ee361318ee2172c7d8d0d184ef20f71b9d4d720533

Observation 9b67ba55-54a6-4052-b681-92f2c2ea78dd · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Large Language Models for Data Annotation and Synthesis: A Survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.014617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.014617Z digest=sha256:83656c8cb23429b92de12328520d472a7ad108262a56471bbb63286cb82e291a

Observation ecf7367b-4479-42d4-a809-59f21da280f5 · outbound

This paper cites Evaluating the Evaluation of Diversity in Natural Language Generation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Evaluating the Evaluation of Diversity in Natural Language Generation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:19.211776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:19.020253Z digest=sha256:6e064f9d84ccf25df7e1e1877c7ee1b1288b0a41437ee4e2d4f7f3511213f776

Observation 1b8f8fe9-b838-442d-9eca-3a0233a5f284 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.026402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.026402Z digest=sha256:71f989a3e60fa4b844aed72417ec6881f2772923dd1e112188f0e21cfdc48b51

Observation e757d0cb-d62c-4142-8bc9-97b74cad6e9e · outbound

This paper cites CodecLM: Aligning Language Models with Tailored Synthetic Data.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning CodecLM: Aligning Language Models with Tailored Synthetic Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.031822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.031822Z digest=sha256:17c2f6ce2355a2d2cf91989a552e4ed0a33ff50a6ae1d052aab5c5f96f778609

Observation 67fb0533-6b1d-479a-83cd-02341d3088a7 · outbound

This paper cites ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.036726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.036726Z digest=sha256:ebb1d1a202bbeb26d29f15a5aaded3b97e968167ce45b43dc413a5330462bbf9

Observation 7d3c46d2-027f-47ac-8757-7cd61c6761b9 · outbound

This paper cites ZeroGen: Efficient Zero-shot Learning via Dataset Generation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning ZeroGen: Efficient Zero-shot Learning via Dataset Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.041283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.041283Z digest=sha256:d4273d23d5d71365aa0531d40510436b5d1a1c6f3de9b4ac34d23a60696d12fa

Observation 19c461a0-09a6-4266-98e1-70345fc5d622 · outbound

This paper cites GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.045722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.045722Z digest=sha256:aeb8c97a764fdfae6186f79d5ac0e8ad0e326e331cc61ca7e89c981186145965

Observation 8804a4a2-3eb3-4378-9620-e07eea7c5218 · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias.Advances in Neural Information Processing Systems, 36, 2024.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Large language model as attributed training data generator: A tale of diversity and bias.Advances in Neural Information Processing Systems, 36, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.654995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:12:19.050725Z digest=sha256:5107ee4584c038826ae6836c4d5f2271fcb0341bee7c81e5693d88a9a7d77318

Observation 82b574e1-7c95-4cdf-9616-d733186d0f80 · outbound

This paper cites How to Synthesize Text Data without Model Collapse?.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning How to Synthesize Text Data without Model Collapse?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.055117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.055117Z digest=sha256:880f501582dbeb62a566be5d7976a506acb6cddb7f5d37d1df544ed52bbe251e

Pith citing papers

Observation f400b560-54ac-4a6e-a286-d40eb55b50fc · inbound

One Joke to Rule them All? On the (Im)possibility of Generalizing Humor cites this paper.

One Joke to Rule them All? On the (Im)possibility of Generalizing Humor What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:39.113934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:53:39.113934Z digest=sha256:786f12b482d30cc468af39e6e8dc35ca4deb9a8bdd0b426018c387c693a9a5ea

Observation f76b29ef-44c0-4475-9f5d-890516a762cf · inbound

Epistemic diversity across language models mitigates knowledge collapse cites this paper.

Epistemic diversity across language models mitigates knowledge collapse What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-03T15:59:04.798474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-03T15:57:18.174038Z digest=sha256:33106085a7ead2c49e2cc012627198e8ece82ad75c81ac2c19ba075ecfceb5b4

Observation 41c68f9a-9d4c-4b38-ae8d-500c2ce2c6f5 · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:11.723762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:11.723762Z digest=sha256:65369834748a361f53439212762cd3bded532f278c425e62594d0b28399f757d