Pith. sign in

REVIEW 4 cited by

Data Distillation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.04272 v2 pith:553RN3Z4 submitted 2023-01-11 cs.LG cs.CVcs.IR

classification cs.LGcs.CVcs.IR
keywords datadistillationapproachesdatasetsresearchsurveytrainingadditionally
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The popularity of deep learning has led to the curation of a vast number of massive and multifarious datasets. Despite having close-to-human performance on individual tasks, training parameter-hungry models on large datasets poses multi-faceted problems such as (a) high model-training time; (b) slow research iteration; and (c) poor eco-sustainability. As an alternative, data distillation approaches aim to synthesize terse data summaries, which can serve as effective drop-in replacements of the original dataset for scenarios like model training, inference, architecture search, etc. In this survey, we present a formal framework for data distillation, along with providing a detailed taxonomy of existing approaches. Additionally, we cover data distillation approaches for different data modalities, namely images, graphs, and user-item interactions (recommender systems), while also identifying current challenges and future research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Clustering of Seafloor Imagery for Interpretation during Long-Term AUV Operations

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An online clustering framework with dynamic cluster splitting/merging and fixed-size representative sampling achieves about 0.68 average F1 on three seafloor image datasets with bounded runtime.

  2. AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ADADEDUP adaptively prunes object detection datasets by combining semantic clustering with proxy-model loss feedback, matching full-data mAP at 20% pruning.

  3. Simple yet Effective Graph Distillation via Clustering

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ClustGDD distills large graphs by clustering node embeddings and refining synthetic attributes, achieving state-of-the-art node classification accuracy at orders of magnitude lower time cost.

  4. MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    MiniLongBench, a 237-sample compression of LongBench, is claimed to reproduce model rankings with a 0.97 Spearman correlation at 4.5% of the evaluation cost.

Pith tools