REVIEW 8 cited by
Datasets: A Community Library for Natural Language Processing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The scale, variety, and quantity of publicly-available NLP datasets has grown rapidly as researchers propose new tasks, larger models, and novel benchmarks. Datasets is a community library for contemporary NLP designed to support this ecosystem. Datasets aims to standardize end-user interfaces, versioning, and documentation, while providing a lightweight front-end that behaves similarly for small datasets as for internet-scale corpora. The design of the library incorporates a distributed, community-driven approach to adding datasets and documenting usage. After a year of development, the library now includes more than 650 unique datasets, has more than 250 contributors, and has helped support a variety of novel cross-dataset research projects and shared tasks. The library is available at https://github.com/huggingface/datasets.
Forward citations
Cited by 8 Pith papers
-
LLM Performance for Code Generation on Noisy Tasks
LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.
-
In-Context Learning (and Unlearning) of Length Biases
LLMs learn length biases from the examples in their prompt, and rebalancing those examples can offset a length bias created by finetuning.
-
Stylometry recognizes human and LLM-generated texts in short samples
Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...
-
ATGen: A Framework for Active Text Generation
The paper presents ATGen, a unified open-source framework for active learning in text generation, with benchmarks showing smart example selection reduces annotation effort and LLM API costs.
-
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
An iterative hybrid pruning method selects which PEFT modules to keep at each transformer layer, matching or improving fixed PEFT baselines on GLUE at 1% trainable parameters.
-
Beyond Simple Concatenation: Fairly Assessing PLM Architectures for Multi-Chain Protein-Protein Interactions Prediction
A systematic benchmark and architecture comparison shows that hierarchical pooling and pooled cross-attention usually beat concatenation for PLM-based protein-protein binding affinity prediction, although statistical ...
-
Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes
Statistical classifiers built on LLM activation norms and coordinates match or beat trained MLP heads on coarse intent routing and resist camouflage better, while MLPs win on fine-grained subfield distinctions.
-
Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs
ID-SPAM generates input-dependent soft prompts with a self-attention mechanism and a two-layer MLP, and shows modest gains over several soft-prompt baselines on NLU tasks.
Discussion (0). Continue with ORCID to comment.