REVIEW 5 cited by
Batch Active Learning Using Determinantal Point Processes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Data collection and labeling is one of the main challenges in employing machine learning algorithms in a variety of real-world applications with limited data. While active learning methods attempt to tackle this issue by labeling only the data samples that give high information, they generally suffer from large computational costs and are impractical in settings where data can be collected in parallel. Batch active learning methods attempt to overcome this computational burden by querying batches of samples at a time. To avoid redundancy between samples, previous works rely on some ad hoc combination of sample quality and diversity. In this paper, we present a new principled batch active learning method using Determinantal Point Processes, a repulsive point process that enables generating diverse batches of samples. We develop tractable algorithms to approximate the mode of a DPP distribution, and provide theoretical guarantees on the degree of approximation. We further demonstrate that an iterative greedy method for DPP maximization, which has lower computational costs but worse theoretical guarantees, still gives competitive results for batch active learning. Our experiments show the value of our methods on several datasets against state-of-the-art baselines.
Forward citations
Cited by 5 Pith papers
-
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
SeDi-Instruct generates instruction data by relaxing duplicate filtering, sampling cluster-balanced batches, and replacing low-scoring seeds with instructions from high-gradient-norm batches.
-
Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning
BADGE-Greedy-DPP selects active-learning batches by greedily maximizing the volume of gradient embeddings, improving rare-call-type discovery in long-tailed frame-level bioacoustic classification.
-
Variational Proximal Policy Optimization
VP2O maps PPO to SVGD in a MoE architecture using functional kernels and expert orthogonalization, claiming +179 ELO on Codeforces and 32% token reduction on AIME for a 33B/4B model.
-
Active learning for photonic crystals
Analytic LL-BNN active learning achieves up to 2.7x reduction in training data for band gap prediction in 2D two-tone photonic crystals while maintaining accuracy.
-
Determinantal point process sampling for bioacoustic active learning
CARE-DPP combines class-balanced uncertainty, annealed embedding novelty, and DPP-based batch diversification for bioacoustic active learning, achieving 0.50 mean AULC versus 0.46 for CoreSet.
Discussion (0). Continue with ORCID to comment.