Pith. sign in

REVIEW 8 cited by

Datasets: A Community Library for Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.02846 v1 pith:RG2SMUV3 submitted 2021-09-07 cs.CL

classification cs.CL
keywords datasetslibrarycommunitynovelsupporttasksvarietyadding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The scale, variety, and quantity of publicly-available NLP datasets has grown rapidly as researchers propose new tasks, larger models, and novel benchmarks. Datasets is a community library for contemporary NLP designed to support this ecosystem. Datasets aims to standardize end-user interfaces, versioning, and documentation, while providing a lightweight front-end that behaves similarly for small datasets as for internet-scale corpora. The design of the library incorporates a distributed, community-driven approach to adding datasets and documenting usage. After a year of development, the library now includes more than 650 unique datasets, has more than 250 contributors, and has helped support a variety of novel cross-dataset research projects and shared tasks. The library is available at https://github.com/huggingface/datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM Performance for Code Generation on Noisy Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.

  2. In-Context Learning (and Unlearning) of Length Biases

    cs.CL 2025-02 conditional novelty 6.0 of 10

    LLMs learn length biases from the examples in their prompt, and rebalancing those examples can offset a length bias created by finetuning.

  3. Stylometry recognizes human and LLM-generated texts in short samples

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...

  4. ATGen: A Framework for Active Text Generation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The paper presents ATGen, a unified open-source framework for active learning in text generation, with benchmarks showing smart example selection reduces annotation effort and LLM API costs.

  5. PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs

    cs.LG 2025-06 conditional novelty 5.0 of 10

    An iterative hybrid pruning method selects which PEFT modules to keep at each transformer layer, matching or improving fixed PEFT baselines on GLUE at 1% trainable parameters.

  6. Beyond Simple Concatenation: Fairly Assessing PLM Architectures for Multi-Chain Protein-Protein Interactions Prediction

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A systematic benchmark and architecture comparison shows that hierarchical pooling and pooled cross-attention usually beat concatenation for PLM-based protein-protein binding affinity prediction, although statistical ...

  7. Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes

    cs.CL 2026-08 conditional novelty 4.0 of 10

    Statistical classifiers built on LLM activation norms and coordinates match or beat trained MLP heads on coarse intent routing and resist camouflage better, while MLPs win on fine-grained subfield distinctions.

  8. Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs

    cs.CL 2025-06 conditional novelty 4.0 of 10

    ID-SPAM generates input-dependent soft prompts with a self-attention mechanism and a two-layer MLP, and shows modest gains over several soft-prompt baselines on NLU tasks.

Pith tools