Pith. sign in

REVIEW 1 cited by

Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07625 v7 pith:B563JTAY submitted 2024-02-12 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords dataautomathtextautodsmathematicalselectionautonomousavailableclassifiers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Autonomous Data Selection (AutoDS), a method that leverages base language models themselves as zero-shot "generative classifiers" to automatically curate high-quality mathematical texts. Unlike prior approaches that require human annotations or training a dedicated data filter, AutoDS relies solely on a model's logits to determine whether a given passage is mathematically informative and educational. By integrating AutoDS into a continual pretraining pipeline, we substantially boost downstream performance on challenging math benchmarks (MATH, GSM8K, and BBH) while using far fewer tokens than previous methods. Empirically, our approach achieves roughly a twofold improvement in pretraining token efficiency over strong baselines, underscoring the potential of self-directed data selection in enhancing mathematical reasoning. We release our curated AutoMathText dataset to facilitate future research in automated domain-specific data curation. The AutoMathText dataset is available at https://huggingface.co/datasets/math-ai/AutoMathText. The code is available at https://github.com/yifanzhang-pro/AutoMathText.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Analyze Feature Flow to Enhance Interpretation and Steering in Language Models

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Cosine similarity between sparse autoencoder features across layers and modules builds flow graphs that explain feature evolution and enable multi-layer steering of language model generation.

Pith tools