Pith. sign in

Well-read students learn better: On the impor- tance of pre-training compact models.arXiv preprint arXiv:1908.08962

12 Pith papers cite this work, alongside 429 external citations. Polarity classification is still indexing.

12 Pith papers citing it
429 external citations · external index

representative citing papers

DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

cs.CL · 2022-10-17 · conditional · novelty 7.0

DiffuSeq adapts diffusion models to conditional sequence-to-sequence text generation and reports performance matching or exceeding strong baselines including pretrained language model systems while generating more diverse outputs.

Geometric and Information Compression of Representations in Deep Learning

cs.LG · 2026-06-19 · unverdicted · novelty 5.0

Experiments show low mutual information does not reliably correspond to geometric compression via class-wise clustering in CEB and dropout networks; the negative nonlinear link can reverse with training changes, suggesting generalization confounds the connection.

PortBERT: Navigating the Depths of Portuguese Language Models

cs.CL · 2026-06-01 · unverdicted · novelty 3.0

PortBERT releases two RoBERTa models for Portuguese that match or beat prior monolingual and multilingual models on translated GLUE/SuperGLUE tasks while reporting training and inference times.

citing papers explorer

Showing 12 of 12 citing papers.