Pith. sign in

REVIEW 2 cited by

SVP-CF: Selection via Proxy for Collaborative Filtering Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.04984 v1 pith:LJUIODYA submitted 2021-07-11 cs.IR

classification cs.IR
keywords samplingperformancealgorithmsdatasetalgorithmdatarelativesvp-cf
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the practical consequences of dataset sampling strategies on the performance of recommendation algorithms. Recommender systems are generally trained and evaluated on samples of larger datasets. Samples are often taken in a naive or ad-hoc fashion: e.g. by sampling a dataset randomly or by selecting users or items with many interactions. As we demonstrate, commonly-used data sampling schemes can have significant consequences on algorithm performance -- masking performance deficiencies in algorithms or altering the relative performance of algorithms, as compared to models trained on the complete dataset. Following this observation, this paper makes the following main contributions: (1) characterizing the effect of sampling on algorithm performance, in terms of algorithm and dataset characteristics (e.g. sparsity characteristics, sequential dynamics, etc.); and (2) designing SVP-CF, which is a data-specific sampling strategy, that aims to preserve the relative performance of models after sampling, and is especially suited to long-tail interaction data. Detailed experiments show that SVP-CF is more accurate than commonly used sampling schemes in retaining the relative ranking of different recommendation algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extending Dataset Pruning to Object Detection: A Variance-based Approach

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A variance-based prediction score using IoU and confidence fluctuations across epochs improves dataset pruning for object detection over several baselines.

  2. FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A greedy token-level Fisher information data selection method that reports improved sample efficiency for GPT-2 supervised fine-tuning on Shakespeare text relative to uniform, density, and AskLLM baselines.

Pith tools