Pith. sign in

REVIEW 3 cited by

DREAM: Efficient Dataset Distillation by Representative Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.14416 v3 pith:GIALSZ3K submitted 2023-02-28 cs.CV

classification cs.CV
keywords matchingdistillationdreamoriginaltextbftrainingdatasetimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dataset distillation aims to synthesize small datasets with little information loss from original large-scale ones for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample synthesis process by matching synthetic images and the original ones regarding gradients, embedding distributions, or training trajectories. Although there are various matching objectives, currently the strategy for selecting original images is limited to naive random sampling. We argue that random sampling overlooks the evenness of the selected sample distribution, which may result in noisy or biased matching targets. Besides, the sample diversity is also not constrained by random sampling. These factors together lead to optimization instability in the distilling process and degrade the training efficiency. Accordingly, we propose a novel matching strategy named as \textbf{D}ataset distillation by \textbf{RE}present\textbf{A}tive \textbf{M}atching (DREAM), where only representative original images are selected for matching. DREAM is able to be easily plugged into popular dataset distillation frameworks and reduce the distilling iterations by more than 8 times without performance drop. Given sufficient training time, DREAM further provides significant improvements and achieves state-of-the-art performances.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unifying Dataset Pruning and Distillation for Efficient Large-scale Compression

    cs.CV 2025-02 conditional novelty 7.0 of 10

    Under a unified ImageNet-1K benchmark, soft labels largely explain the success of large-scale dataset distillation, and a hard-label pipeline that prunes, combines, and augments real images beats prior methods at extr...

  2. Data-to-Model Distillation: Data-Efficient Learning Framework

    cs.CV 2024-11 conditional novelty 6.0 of 10

    D2M distills a dataset's knowledge into the parameters of a pre-trained GAN, enabling flexible, architecture-general synthetic training data with state-of-the-art classification accuracy.

  3. GMem: A Modular Approach for Ultra-Efficient Generative Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GMem conditions diffusion models on a fixed bank of DINOv2 features and reports much lower FID at far fewer epochs than SiT and REPA baselines on ImageNet.

Pith tools