Pith. sign in

REVIEW 1 cited by

On sampling from data with duplicate records

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.10549 v1 pith:YB5NJQVQ submitted 2020-08-24 cs.LG cs.DBcs.IRstat.ML

classification cs.LGcs.DBcs.IRstat.ML
keywords entitiessamplingdatadatabasefrequencyrecordssteptask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data deduplication is the task of detecting records in a database that correspond to the same real-world entity. Our goal is to develop a procedure that samples uniformly from the set of entities present in the database in the presence of duplicates. We accomplish this by a two-stage process. In the first step, we estimate the frequencies of all the entities in the database. In the second step, we use rejection sampling to obtain a (approximately) uniform sample from the set of entities. However, efficiently estimating the frequency of all the entities is a non-trivial task and not attainable in the general case. Hence, we consider various natural properties of the data under which such frequency estimation (and consequently uniform sampling) is possible. Under each of those assumptions, we provide sampling algorithms and give proofs of the complexity (both statistical and computational) of our approach. We complement our study by conducting extensive experiments on both real and synthetic datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DobLIX: A Dual-Objective Learned Index for Log-Structured Merge Trees

    cs.DB 2025-02 conditional novelty 6.0 of 10

    A dual-objective learned index for LSM trees co-optimizes block partitioning and lookup error, with an RL agent tuning parameters, and reports 1.19-2.21x throughput gains in RocksDB.

Pith tools