Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:43:29.483303Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.20465.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:43:29.483303Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bd4b03ec-4b6f-40c5-8b3e-b5487fe65a13 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Fowlkes, Stefano Soatto, and Pietro Perona
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdffe553-588b-46b7-888b-cadc8844152d · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators A Survey of Multimodal Large Language Model from A Data-centric Perspective
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd684c46-25f3-4722-bfef-fec878732057 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Text2sql-flow: A robust sql-aware data augmentation framework for text-to-sql.arXiv preprint arXiv:2511.10192, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a93b4c45-4b63-45cb-be31-147ea1e05d4b · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Data-juicer: A one-stop data processing system for large language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcbe75e0-c1c3-41c6-a1a3-81dff4e03df5 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Dc-bench: Dataset condensation benchmark.Advances in Neural Information Processing Systems, 35:810–822, 2022
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4aa522-9ce7-4455-859b-e95d8a464882 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Enhancing chat language models by scaling high-quality instructional conversations, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 603995aa-e9eb-4334-a2a1-06b30605e8f0 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a22e3b31-e2d7-4070-a10c-45ae2789e857 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941e3dc9-bf3c-4716-a14f-fbaec6ad7448 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators The Vendi Score: A Diversity Evaluation Metric for Machine Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858d25fd-5f67-4e8b-aa18-786e218ef8c7 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Datacomp: In search of the next generation of multimodal datasets
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96272a94-f334-4826-be5b-c904b8f16b8c · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Closing the data loop: Using opendataarena to engineer superior training datasets.arXiv preprint arXiv:2601.09733, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edec8fc0-5fe7-4784-8327-c3dd57a4efee · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8260ba64-d930-43a0-9770-7534233be136 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Textbooks Are All You Need
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6371c9-9507-4444-b855-2c9090b67b95 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Lawyer llama technical report, 2023
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c181c15-a749-47ab-9e2a-19cf6cb0d272 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Mistral 7B
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56135507-05ce-4684-a697-169e33636bf7 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Scaling Laws for Neural Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc07498d-6c21-4743-bacd-6e6d90f45803 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798812e9-500e-42ac-beab-8982f38a73f6 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Datacomp-lm: In search of the next generation of training sets for language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29130112-8ba6-494e-9325-3879df750380 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32307cbf-ed4d-4525-ba4b-b64736018cbd · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Superfiltering: Weak-to-strong data filtering for fast instruction-tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 951c07c5-babc-41dd-9ebd-a3df9f7e8723 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai.arXiv preprint arXiv:2512.16676, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0018cca2-de3d-4cf0-89a8-2fcedc1c0926 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Data preparation for large language models.Journal of Computer Science and Technology, 2026
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b170a5ed-a4c8-4819-8681-4183ffe4a4bf · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e194211-7910-45fc-ab33-883e69dbf12e · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Dataperf: Benchmarks for data-centric ai development
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5fcee22-949f-4808-97ea-f56f89739ed4 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators GPT-4 technical report, 2023
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1908213-b27b-41dd-8382-606338e87975 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Towards tailored recovery of lexical diversity in literary machine translation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0428d130-e587-457f-8628-4a16b3509635 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ccdba98-9f92-4550-a0c8-458feef7a0ab · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators A survey on domain adaptation theory: learning bounds and theoretical guarantees
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae485336-5f26-4516-87d6-4d769b4338dc · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Let’s verify math questions step by step.arXiv preprint arXiv:2505.13903, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b144c66-67dc-4ca7-8326-4f5ae8b151e7 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Reasonmed: A 370k multi-agent generated dataset for advancing medical reasoning,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9efaef-e9bc-4663-b11d-f628591b0d7d · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators LLaMA: Open and Efficient Foundation Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9eb37d7-e4c7-44c3-8ed3-777f9175fe87 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Self-instruct: Aligning language models with self-generated instructions
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4bba88-7a4f-4b5c-b9e0-987ea6a9ec2b · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators QuRating: Selecting High-Quality Data for Training Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d50ac5-3475-4a71-a7fc-109dec7d815d · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators LESS: Selecting Influential Data for Targeted Instruction Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31dd247f-db92-44bf-b5b7-d0f7224e1082 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Wizardlm: Empowering large pre-trained language models to follow complex instructions, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd37b376-0ddf-4eb3-9bed-1cd38b27c9e0 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Logics-stem: Empowering llm reasoning via failure-driven post-training and document knowledge enhancement, 2026
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7bbbe0-d135-48ed-9c48-1b67fee50a50 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Qwen2 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5787fa-93fb-43fc-811f-d053fa1ca382 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Qwen2.5 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1be480d-9b52-4913-8e55-650ef7ec0d5b · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a79cf5e-56d5-45fa-b806-dfec79adf9fb · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Disc-lawllm: Fine-tuning large language models for intelligent legal services, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca1f35df-7d2d-4df3-852a-e25dd474a82b · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Ultramedical: Building specialized generalists in biomedicine, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f808fa4-a4bc-489c-80c0-886038268843 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4720352e-9d9c-4b37-ac62-f800bdecbbd1 · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Google-proof
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27cc783e-2d4e-47a6-837b-b1d93658916a · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a483e706-d926-47cf-b11a-b27d902d965a · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17223647-62dc-4e87-b1fc-eedfb60ea8fc · outbound
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.