Pith. sign in

REVIEW 1 cited by

High-Resource Methodological Bias in Low-Resource Investigations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.07534 v1 pith:V4JYP6P3 submitted 2022-11-14 cs.CL

classification cs.CL
keywords low-resourcedatadatasetsdownhigh-resourceresultssamplinginvestigations
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The central bottleneck for low-resource NLP is typically regarded to be the quantity of accessible data, overlooking the contribution of data quality. This is particularly seen in the development and evaluation of low-resource systems via down sampling of high-resource language data. In this work we investigate the validity of this approach, and we specifically focus on two well-known NLP tasks for our empirical investigations: POS-tagging and machine translation. We show that down sampling from a high-resource language results in datasets with different properties than the low-resource datasets, impacting the model performance for both POS-tagging and machine translation. Based on these results we conclude that naive down sampling of datasets results in a biased view of how well these systems work in a low-resource scenario.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuning open-source LLMs outperforms prompting for POS tagging on medieval Occitan, French, and Spanish, and pooling Romance training data helps the most under-resourced texts.

Pith tools